The open risk since session 2 was "a sustained action sequence could still
break the bitrate", with every clip measured so far being 1.2-1.7 s. Closed by
measurement rather than by sampling clips by hand.
07_motion_survey.py scans a whole stream at 96x72 for the hottest sliding
window of inter-frame difference. On 00223 the spread between the quietest and
hottest sustained 10 s windows is 10.6x, which is the argument for not eyeballing
it. Hottest is t=539.4s, the Singe endgame.
There, with the fixed lam the CLI uses, sasi overshoots 110 -> 129.6 KB/s (+18%)
and scsi 280 -> 373.8 KB/s (+34%). Rate control moves from "insurance, not a
fix" to required, and is promoted above the full-disc survey. The bus is not
broken -- 381.6 KB/s still fits the 488 KB/s figure -- so FINDINGS 21 survives,
at 78% of the pipe instead of a comfortable margin.
Three further corrections fall out:
- The two largest streams on the disc are bonus material. 00216 is the feature
with a burned-in commentary PiP; 00215 is the commentary. 00223 is the clean
9.4 min. A size-ranked survey would have encoded live action.
- On hard content the 256-colour scene palette (31.33 dB) binds well before the
X68000 display (40.81 dB); scsi is already within 0.51 dB of it.
- FINDINGS 24.5's architecture question resolves to "both paths, chosen per
frame": 30-53% of frames sit above the 70% crossover. Picking per frame costs
a median 37.0% of the frame budget and caps at 53.6%. Reporting for this is
wired into encode.py, which previously only printed a mean over all frames --
the one statistic that cannot answer a per-frame question.
extract.py takes optional start/dur; 08_mode_map.py renders source | decoded |
block-mode map to .webm.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Answers session 1's critical-path question. Flat 4x4 VQ at k=256 was prototyped
and REJECTED by eye: Dirk's face disintegrates and ink outlines break into
4-pixel stair-steps. The 256-colour palettised frame is excellent, so the
palette was never the problem -- block VQ was.
Replaced it with a Cinepak-style hybrid: each 4x4 block is SKIP, one 4x4
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
rate-distortion. The RAW escape makes lam=0 pixel-exact (measured 0.00 dB loss),
so the quality knob spans lossless to heavily-compressed in one bitstream.
Per the user's decision, ships TWO quality profiles from that one codec, one
decoder and one bitstream -- only the rate knob differs:
sasi 45 KB/s lam=300 34.8 dB stock 10MHz ACE/EXPERT
scsi 75 KB/s lam=100 35.9 dB Super/XVI or CZ-6BS1
Three corrections to earlier numbers:
1. Session 1's "183 KB/s at 12fps" was a bad extrapolation. Halving the
framerate does not halve the bitrate -- decimation roughly doubles the
per-frame delta. Re-measured directly: 340 KB/s for session 1's own RLE,
247 KB/s for changed-spans+deflate. The lossless floor is 319 MB.
2. A FOURTH false-good result, same family as the three in FINDINGS 4:
k=1024 codebooks appeared to buy +2.4 dB free, because the rate model
charged 1 byte for a 10-bit index. Charging the true cost reverses the
verdict -- k=256 wins at every matched bitrate, and by 5 dB at the low end
where the SASI profile lives. k=256 ships.
3. Stream inventory: the ~3-5MB clips are 1.2-1.7s, not ~60s, and some 60s
streams are menus, not content. Any survey must classify before averaging.
Also cleared both candidate sources for the game-logic layer: the SNES project
is MIT and DirkSimple is zlib, so the arcade scene graph can be imported and
the two transcriptions diffed against each other.
Encoder is working end-to-end: extract.py -> vq/vq_hybrid/ratectl -> encode.py,
emitting a big-endian DLX1 container the 68000 can parse with plain moves.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6