Files
Dragon-s-Lair-X68k/docs/STATUS.md
T
prosolis 3265bf2740 Handoff: reconcile docs with the verified display path
Session 3 summary in STATUS.md, plus the things a cold start needs.

- Reproduce section for the display result, verified cold from the Blu-ray at
  end of session: extract -> prep -> MAME -> verify, exact match, 38.88 dB.
  The frames are not in the repo and the old ones lived in /tmp, so the chain
  starts from extract.py rather than assuming a scratch directory survives.
- tools/bench/verify_frame.py turns FINDINGS 22 into a regression check. It is
  deliberately an exact test rather than a PSNR threshold, since the whole
  point of that section is that the render is bit-for-bit predictable. It
  prints the three registers to check when it fails.
- Recorded where the MAME source now lives, and why to read it first: six
  register-poking attempts failed against a gate that one grep found.
- Split the CRTC mode table out as its own next step. It is the prerequisite
  for the decoder skeleton and the smallest well-defined task available, with
  an explicit warning not to write the timing values from memory.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 13:16:04 -07:00

16 KiB

Status & next-session handoff — end of session 3 (2026-08-23)

Decisions locked

decision value why
Target CPU 68000 @ 10MHz (stock) hardest honest constraint
Display mode 256 colors, 256x192 in 256x256 CRTC mode every mode is 1 word-access/pixel, so 256c is free vs 16c
Double buffer none — page 1 sacrificed enables movem.l 24px bursts; delta coding needs a RAM reference frame anyway
Codec hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion flat 4x4 VQ was measured and rejected — see FINDINGS 9-10
Quality modes two: sasi and scsi (USER DECISION, session 2) one codec, one decoder, one bitstream; only lam differs
Framerate 12 fps, explicit decimation source has zero duplicate frames; no free "twos" win
Emulator MAME 0.277 x68000 accurate enough that measured cycles mean something
SNES project reuse MIT — cleared data/events/ scene graph is reusable with attribution

The SASI/SCSI question is RESOLVED

Session 1 left "which machine do we target" open. The user's answer: ship both, as two quality profiles. This is now implemented rather than hypothetical — the bitrate ceiling is a build parameter in tools/encoder/ratectl.py:

profile target lam quality (00020 / 00146) machine
sasi 110 KB/s 60 36.9 / 29.6 dB stock 10MHz ACE/EXPERT
scsi 280 KB/s 10 39.4 / 32.3 dB Super/XVI, or CZ-6BS1 board

Sized against the user's working figure of 4 Mbps = 488 KB/s sustained, on SD-backed SCSI (BlueSCSI / SCSI2SD) — so that rate is a bus-limited constant, not an average over seek latency.

Both profiles fit with room. Ring-buffer simulation on the real per-frame sizes gives zero required prefill for every scene at both profiles: the fill delivers 40.69 KB per frame time and only one measured frame (42.10 KB) exceeds that, recovered by the next. A 256 KB buffer carries ~1 s of stall tolerance, far more than an SD-backed seek needs. FINDINGS 21.

An earlier warning here said scsi did not fit because a frame peaked at 96.4% of the pipe. That compared instantaneous demand to a sustained rate as if they had to match frame-by-frame; with a buffer the test is cumulative, and it passes.

scsi is now within 0.5 dB of the palette ceiling on 00020. These were initially set at 45 / 75 KB/s, which was 12% / 7% bus utilisation — read off the RD curve rather than derived from the hardware. See FINDINGS 17.

Codebooks are k=256 with 1-byte indices in both profiles. k=1024 was measured and rejected — see FINDINGS 14, it was a false-good result from a rate model that undercharged the index. Do not ship past lam~800; FINDINGS 15 has the cliff.

Because of the RAW escape mode, lam=0 is pixel-exact against the palettised frame (measured 0.00 dB loss). The profiles are two points on one continuous rate-distortion curve, not two codecs.


What session 3 settled

  1. The display path works and is verified end to end. First real frame on an emulated X68000 screen: docs/images/x68k_first_frame_compare.png. Full write-up in FINDINGS 22. Everything before this session was Python-side or a headless -video none run, which cannot snapshot at all.
  2. The render is pixel-exact, not merely close. With monitor contrast at 15, all 256 palette entries render exactly as GGGGGRRRRRBBBBBI + pal6bit predicts. That exactness is the regression test — see tools/bench/verify_frame.py, which exits non-zero if it ever drifts.
  3. Three hardware facts that were previously assumed are now confirmed from MAME 0.277 source, not folklore: the palette word format, the 1024-byte GVRAM line stride, and the 256-colour page aliasing in HARDWARE.md. All three were already written down correctly; they are now cited.
  4. A new quality ceiling was measured — the 15-bit+I palette alone costs 38.88 dB, the same order as the scsi profile's own codec error. FINDINGS 22.4. This bounds how much further scsi is worth raising.
  5. Two shell traps that wedged session 2's background jobs are documented in the working-setup section below. They cost ~1.5 h of wall clock and a wedged CPU core, and one of them was hit again this session.

What session 2 settled

  1. The critical-path question is answered. "Does VQ soften Bluth's linework unacceptably?" — flat 4x4 VQ: yes, badly. The hybrid (SKIP/V1/V4/RAW): no. Verified by eye, not just PSNR. See docs/FINDINGS.md 9-11 and the two images in docs/images/. Both profiles use k=256; see item 2b.
  2. Session 1's 12fps bitrate was wrong (183 KB/s claimed, 340 KB/s measured). Halving the framerate does not halve the bitrate. FINDINGS 8. 2b. A fourth false-good result was produced and caught this session — k=1024 codebooks looked like a +2.4 dB free win because the rate model charged 1 byte for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
  3. The 256-colour palettised frame is the real quality ceiling and it looks excellent. Judge the codec against that, not against 1080p.
  4. Encoder exists and produces a real bitstream: tools/encoder/.

Encoder — working

python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
python3 tools/encoder/encode.py  /tmp/fr_00020 out.dlx --profile sasi --preview p.png
file role
extract.py .m2ts -> 256x192 PNGs, 12fps, spatial-only denoise
vq.py palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR
vq_hybrid.py the codec: 4 block modes + lagrangian mode decision
ratectl.py SASI/SCSI profiles, leaky-bucket rate control
encode.py CLI + DLX1 container writer

DLX1 container layout is documented in the encode.py docstring. All multi-byte fields are big-endian so the 68000 reads them with a plain move.

Known encoder gaps

  • Rate control is written but not yet wired into encode.py — the CLI uses a fixed lam from the profile. ratectl.encode_rate_controlled() exists and builds a lam-ladder per frame; it needs hooking up and validating.
  • Payload is deliberately NOT entropy-coded — deflate decode does not fit in the 68000's frame budget (FINDINGS 17.2). Do not "optimise" this later.
  • Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
  • _paint is a Python per-block loop — fine for prototyping, slow for a full disc encode. Vectorise before the 224-stream run.

Working setup (unchanged from session 1, re-verified)

MAME ROMs~/mame/roms/x68000.zip. Must pass -bios ipl10.

mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3

Assemblertools/vasm/vasmm68k_mot -Fbin -o out.bin in.s

Blu-rayudisksctl loop-setup -r -f DRAGONS_LAIR.iso -> /media/reala-misaki/BDROM (still mounted as of end of session 2).

MAME Lua harnesstools/bench/*.lua, working. Three gotchas (retain the notifier subscription in a global; the stack register is SP not A7; autoboot_script fires at PC=0 before boot) are documented in FINDINGS.

Two shell traps, both hit again this session:

  • piping MAME (or any long job) through grep block-buffers — write to a file.
  • pkill -f <pattern> matches your own shell and kills it (exit 144). Use pkill -x or kill by PID.
  • until ! pgrep -f foo.py; do sleep; done watcher loops never exit. The watching shell's own command line contains the string foo.py, so pgrep -f matches the watcher itself and the loop spins forever. Session 2 left 11 of these wedged for over an hour. Wait on the PID (while kill -0 $PID) or on a sentinel file the job touches when it finishes -- never on a -f name match.
  • timeout N mame ... does not kill MAME. MAME catches SIGTERM and, with an autoboot script blocked waiting on a flag that never arrives, never reaches its shutdown path. timeout without -k then waits forever while MAME burns a full core at -nothrottle. Always timeout -k 5 N.

Disk throughput benchmark — still blocked, no longer gating

IOCS _B_READ returns -1 uniformly. Full diagnosis and the four untested hypotheses are in session 1's notes (git history of this file, commit 65112b9); the ordered plan for retrying is in docs/BENCHMARK.md.

Status changed twice this session — read this rather than the git history. It was briefly promoted to critical-path while the working bandwidth figure was misread as 4 MB/s. With the correct figure (4 Mbps = 488 KB/s) and the ring-buffer simulation showing zero required prefill for both profiles (FINDINGS 21), the design no longer hangs on it. Pixel-exact on SCSI is not available at 4 Mbps — it needs 92-97% of the pipe — so there is no longer a "measure it and maybe ship transparent" decision waiting.

What the benchmark is still worth doing for:

  • Confirming the 4 Mbps figure. It is user-supplied and its provenance is not recorded. Every profile hangs off it.
  • Confirming DMA is actually used. If transfers fall back to PIO the CPU cost rises far above the ~12-15% cycle-steal estimate and CPU becomes the binding constraint. This is the worst plausible outcome and the cheapest to check — do it first.

Do not try to get the bandwidth number out of MAME. Its SCSI/SASI devices are functional models, not timing-accurate; a KB/s figure from MAME measures the emulator's scheduler. docs/BENCHMARK.md covers the three-tier approach (MAME validates the path, derivation bounds it, real hardware settles it).

Display path — WORKING, verified end to end (session 3)

The first real frame is on screen: docs/images/x68k_first_frame_compare.png. Full write-up in FINDINGS 22. Harness: tools/bench/show_frame.lua + tools/bench/prep_frame.py.

Three facts the player MUST honour, none of which were guessable:

what where value
Un-hide the graphics layer CRTC R20 $E80028 clear bit 11 ("G-VRAM set to buffer"); IPL leaves 0x0B16
Colour setup (256c) CRTC R20 bits 9-8 0x0100 -> R20 = 0x0116
Monitor contrast $E8E001 bits 3-0 IPL leaves 14; write 15 or everything renders 7% dark

Bit 11 is the one that cost the most time: GVRAM writes land and read back correctly while the layer is invisible, so the video controller looks guilty and is not. Contrast 0 blanks the screen — free fade-to-black for transitions.

Palette format is now confirmed from MAME source, not assumed: GGGGGRRRRRBBBBBI (G 15:11, R 10:6, B 5:1, shared LSB I), expanded as pal6bit((field<<1)|I). With contrast at 15 the render is pixel-exact.

New ceiling: the 15-bit+I palette alone costs 38.88 dB against the 24-bit palettised source — the same order as the scsi profile's own codec error (39.4 dB). scsi is close to display-transparent on real hardware. See FINDINGS 22.4 before considering raising quality further.

Snapshot recipe that works (-video none CANNOT snapshot):

SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
  -sound none -nothrottle -plugins -autoboot_script <script>.lua \
  -snapshot_directory ./snap -snapview native -seconds_to_run 6

-snapview native drops MAME's LED artwork and gives a clean 768x512 screen.

Next steps, in priority order

  1. Full-disc survey. Only 4 clips of 1.2-1.7 s out of 224 streams have been measured, and 00146 already runs 23% hotter than 00020. A sustained action sequence is the one thing that could still break the bitrate. Classify menu vs content first (FINDINGS 13) or the averages are diluted by static menus. Vectorise _paint before this run — it is a Python per-block loop.

  2. 68000 decoder skeleton. Parse DLX1, expand codebooks to word-per-pixel, blit SKIP/V1/V4/RAW. Measure real cycles with the existing MAME Lua harness. Now unblocked — the display path is verified (FINDINGS 22) and tools/bench/show_frame.lua gives a known-good reference image to diff the 68000's output against. Validates the 38% full-frame blit estimate that the whole CPU budget rests on. Still needs a real CRTC mode table for 256x256; the harness deliberately borrows the IPL's timing and invents nothing. 2a. CRTC mode table for 256x192-in-256x256. Prerequisite for (2) and the smallest well-defined unit of work available right now. Needs real R00-R08 timing values. Do not write these from memory — session 3 lost time to exactly that failure mode on the video registers. Derive them from the CRTC dividers in x68k_crtc.cpp (m_reg[20] & 0x1f selects the dot-clock divisor; the IPL's 0x16 gives /2 off the 69MHz clock), or lift a known-good set from a real X68000 title and verify by snapshot. The harness makes this cheap to iterate: change values, snapshot, look.

  3. Wire rate control into encode.py. No longer a blocker (FINDINGS 21), but it is what gives a deterministic ceiling over content not yet measured, which was the original reason for choosing VQ. Insurance, not a fix. Pairs with (1).

  4. Confirm DMA vs PIO in MAME (see the benchmark section above) — cheap, and the only thing that could still move CPU into the binding position.

  5. Resolve the framing question (FINDINGS 12: crop vs squash vs wide). Needs an eyeball against arcade reference, not a measurement.

  6. Import the scene graph. SNES project data/events/ (MIT, cleared), cross-checked against DirkSimple (zlib) which transcribed the same data independently — diff them to catch transcription errors before committing any of it to 68000 tables.

  7. ADPCM audio. MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in ratectl.py, not yet extracted or encoded.

Explicitly abandoned — do not re-propose

  • Entropy-code the payload. Deflate decode is ~216% of the frame budget on a 68000; LZ4 is ~54% with no room beside a 38% blit (FINDINGS 17.2). All bitrates are raw payload. This also demotes the "247 KB/s lossless" figure in FINDINGS 8 to a compression upper bound, not a shippable design.
  • k=1024 codebooks. False-good result from a rate model that charged 1 byte for a 10-bit index (FINDINGS 14). k=256 wins at every matched bitrate.
  • Flat 4x4 VQ. Rejected by eye (FINDINGS 9).

Not yet started

  • Any 68000 player code. src/player/ is still empty. The display path is proven, but proven from Lua — no 68000 instruction has yet drawn a pixel.
  • A real CRTC mode table. The harness deliberately borrows the IPL's 768x512 text timing and invents no CRTC values, which is why the frame repeats at x=512 (FINDINGS 22.5). A 256x256 mode needs real R00-R08 values, and those must be derived or measured, NOT recalled from memory — see the note below.
  • ADPCM audio extraction/encoding
  • Disk image packaging
  • Game logic (scene branching, input windows, death clips)

Reproducing the display result

python3 tools/encoder/extract.py 00020 tmp/fr_00020 12 crop
python3 tools/bench/prep_frame.py tmp/fr_00020 tmp/frame.bin 0
mkdir -p tmp/snap_verify && cd tmp && SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 \
  -video soft -window -sound none -nothrottle -plugins \
  -autoboot_script ../tools/bench/show_frame.lua \
  -snapshot_directory ./snap_verify -snapview native -seconds_to_run 6
cd .. && python3 tools/bench/verify_frame.py

Verified cold from the Blu-ray at end of session 3: exact match, 38.88 dB.

tmp/ is gitignored scratch. The frames are NOT in the repo — regenerate them with extract.py; the earlier ones lived in /tmp and do not survive a reboot.

Reference material on this box (not in the repo)

  • MAME 0.277 source: ~/src/mame-mame0277/ (tarball ~/src/mame0277.tar.gz). Downloaded this session to settle the graphics-layer question. The files that matter are src/mame/sharp/x68k_v.cpp, x68k_crtc.cpp, x68k_crtc.h, x68k.cpp. Read these before theorising about X68000 video behaviour — six register-poking attempts failed against a gate that one grep found.
  • Blu-ray mounted at /media/reala-misaki/BDROM via udisksctl loop-setup -r -f DRAGONS_LAIR.iso.