FINDINGS 26 stopped the session-5 rate controller before it shipped: it built a
lam-ladder of independent whole-sequence encodes and picked frames off it, so
SKIP blocks referenced reconstructions the decoder never saw -- 111 of 120
frames drifted. The fix is the structural one 26.1 said it had to be.
vq_hybrid is now frame-drivable -- frame_ctx / decide / paint -- and encode() is
a thin loop over it. Rate control drives the same three calls, bisects lam per
frame under the leaky bucket, and feeds back the frame it actually emitted. The
desync has no way to occur, and 09_ratectl_drift.py goes 111/120 -> 0/120. That
test is now part of check.sh, which is ~2 min rather than ~40 s.
Both overshoots on the worst sustained window are closed for under 1 dB, totals
including audio: sasi 137.4 -> 109.5 KB/s (-0.60 dB), scsi 381.6 -> 280.0 KB/s
(-0.91 dB). Zero frames hit the lam=800 cliff, so nothing was destroyed to get
there. Rate control also makes the display path cheaper -- scsi's median drops
53.6% -> 47.1% -- because raising lam moves blocks to SKIP and V1.
Two knobs measured rather than guessed. --rc-floor is worth 0.00 dB on that
window and defaults to the profile lam, so rate control cannot regress content
that already fits. --prefill defaults to 0 and is documented as a trap: it buys
a permission to overshoot of exactly bucket/nframes, and on a 14-frame clip it
disables the controller outright.
FINDINGS 26.5 was wrong in both halves and 27.6 records it. _paint was not the
bottleneck (14% of a frame, though vectorising it was still right at 17.1x) and
the ladder was never "minutes" -- those were k-means in build(). What makes
per-frame rate control affordable is that VQ.assign depends on neither lam nor
prev, so it is cached one frame deep: a 12-step search over 120 frames costs
0.31 s against 49.1 s.
Also caught: fixed-lam sasi was already 5% over target on 00020, the clip
everyone called easy. Nothing noticed because the profile table quotes PSNR and
not bitrate.
check.sh: ALL GREEN.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
First 68000 instructions in this project to draw a pixel. Everything before
this was GVRAM filled from Lua, which costs zero 68000 cycles, so the blit
figure the whole CPU budget rests on had never been validated.
Four variants of a full-frame 256x192 paint, timed in MAME and each also
hand-derived from the MC68000 timing tables beforehand; the two agree to
0.006-0.43%, which is what makes the result trustworthy after this project's
history of false-good measurements.
V1 movem.l blit from a word-expanded RAM frame 446,286 cyc 53.6%
V2 naive move.b/move.w per pixel 1,284,174 cyc 154.1%
V3 write-only floor, no source read 225,789 cyc 27.1%
V4 same writes in 4x4 block order 637,971 cyc 76.6%
Scope: MAME's gvram_w/gvram_r carry no timing at all, so these are instruction
cycles against zero-wait-state memory -- a floor, not a hardware prediction.
V1's output snapshots pixel-exact through verify_frame256.py, closing
FINDINGS 23.5. The V1/V3 gap shows reading the source frame is exactly half
the cost, which makes the architecture question live: decode-direct-to-GVRAM
needs no RAM reference frame and scales with the non-SKIP block fraction,
crossing compose-then-blit at 70% of blocks changed. That fraction is now the
top priority and is already a by-product of vq_hybrid.py's mode decision.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Prepares session 4 for handoff. No new measurements; this reconciles the docs
with what session 4 changed and makes the next session's entry point cheaper.
- tools/bench/check.sh re-runs both display regression tests from the Blu-ray in
~40 s and prints ALL GREEN. Verified green cold, after wiping tmp/ and
re-extracting. STATUS and README both open with it, because everything
downstream assumes the display path is pixel-exact and nothing previously
checked that in one command.
- Reconciled the figures session 4 invalidated. Session 3's 38.88 dB ceiling is
struck through in STATUS with a pointer to 40.81; the "three facts the player
must honour" table no longer quotes R20 = 0x0116, which was the 768-wide IPL
timing and would have been copied into the player as if it were the shipping
value. crtc_mode.lua is now named as the single source of truth for CRTC
registers, in both STATUS and README.
The 38.88 dB in the session-3 reproduce section is left alone and annotated
instead: it is correct for that test, which still packs I = 1. The two numbers
disagree for a reason and a reader should be able to see which is which.
- Next-step 2 now leads with something smaller than "write the decoder": a dumb
full-frame RAM->GVRAM blit in 68000 code, timed. That number alone confirms or
kills the 38% estimate, and needs no bitstream, codebooks, or DLX1 parsing.
The reference image and its checker already exist.
- Parked the user's Cliff Hanger / Lupin III follow-on in STATUS so it is not
lost and not mistaken for scheduled work. Cheaper than this project on every
axis except media prep, which is where it would actually stall.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Session 3 left the harness on the IPL's 768x512 text timing because no CRTC
values had been derived and guessing them was the failure mode to avoid. This
derives them from MAME 0.277's divisor ladder instead, and the derivation is
self-checking: the 256-wide mode runs at div 6 against the 768 mode's div 2, so
htotal is exactly 1104/3 = 368 dots and every horizontal register divides by
three with no remainder. Only the blanking split rounds. Verified by snapshot:
native 256x512, active area pixel-exact, x=512 wrap gone.
Two things fell out that change numbers elsewhere:
- The palette's shared LSB I must be chosen per entry, not hardcoded to 1.
Doing so lifts the display ceiling from 38.85 to 40.81 dB and is the only way
to reach true black at all, since pal6bit(1) = 4. 102 of 256 entries want
I = 0, so this is not a corner case. Supersedes FINDINGS 22.4; scsi has ~2 dB
more headroom than that section claimed. The encoder does not do this yet.
- Letterboxing costs a palette entry: GVRAM cleared to zero shows entry 0, and
a free mediancut palette puts a real image colour there. 255 colours plus a
reserved black, via prep_frame.py --reserve-black.
MAME's graphics double-scan is phase-shifted one raster line (it halves the
absolute scanline and vbegin is odd), which produced a false failure before it
was understood; the regression test now asserts the shifted pairing explicitly.
Still Lua-side. No 68000 instruction has drawn a pixel; the 38% blit estimate
remains unvalidated. What this buys is a defined geometry for the decoder to
write into: 256 words per row, 1024-byte stride, rows 32..223.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Session 3 summary in STATUS.md, plus the things a cold start needs.
- Reproduce section for the display result, verified cold from the Blu-ray at
end of session: extract -> prep -> MAME -> verify, exact match, 38.88 dB.
The frames are not in the repo and the old ones lived in /tmp, so the chain
starts from extract.py rather than assuming a scratch directory survives.
- tools/bench/verify_frame.py turns FINDINGS 22 into a regression check. It is
deliberately an exact test rather than a PSNR threshold, since the whole
point of that section is that the render is bit-for-bit predictable. It
prints the three registers to check when it fails.
- Recorded where the MAME source now lives, and why to read it first: six
register-poking attempts failed against a gate that one grep found.
- Split the CRTC mode table out as its own next step. It is the prerequisite
for the decoder skeleton and the smallest well-defined task available, with
an explicit warning not to write the timing values from memory.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
First pixels on an actual X68000 screen. Everything up to now was Python-side
or a headless -video none run, which cannot snapshot at all.
The blocker was not the video controller. The IPL leaves CRTC R20 = 0x0B16,
and bit 11 is "G-VRAM set to buffer", which makes MAME's draw_gfx() return
early. GVRAM writes still land and read back correctly while the layer is
invisible, so six attempts at $E82400/$E82500/$E82600 all rendered black with
every register holding the value I intended.
Two more facts, both confirmed against MAME 0.277 source rather than assumed:
- $E8E001 monitor contrast is left at 14 by the IPL, scaling all output to
93.3%. The player must set it to 15. Contrast 0 blanks the screen, which is
a free fade-to-black for scene transitions.
- The palette word is GGGGGRRRRRBBBBBI with a shared LSB, expanded as
pal6bit((field<<1)|I). With contrast at 15 the render is pixel-exact, not
merely close, which also confirms the 1024-byte GVRAM line stride.
That exactness gives a new quality ceiling: the 15-bit+I palette alone costs
38.88 dB against the 24-bit palettised source, the same order as the scsi
profile's own codec error. scsi is close to display-transparent on hardware,
which bounds how much further it is worth raising.
Unblocks next step 2, the 68000 decoder skeleton, which now has a known-good
reference image to diff against.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Verified GVRAM is one word-access per pixel in ALL color modes; chose 256-color
256x192 with movem.l bursts (page 1 sacrificed as double-buffer).
Measured 8 scenes from the Blu-ray source: blit costs under 8% of the 12fps
cycle budget, so I/O is the bottleneck, not CPU. Naive delta+RLE reaches only
3.2:1 (365 KB/s, 470MB) -> decision to use 4x4 vector quantization (~30 KB/s).
"Shot on twos" assumption failed: the transfer has zero duplicate frames, so
12fps requires explicit decimation.
Documents three false measurement results and their root causes (per-frame
Floyd-Steinberg dithering, temporal denoise, exact-match dedupe on noisy source).
MAME Lua injection harness works and is reusable for cycle-cost measurement;
the IOCS _B_READ disk benchmark is blocked returning -1.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6