Files
Dragon-s-Lair-X68k/docs/STATUS.md
T
prosolis e4062ed294 Session 2: hybrid VQ codec, two quality profiles, three corrections
Answers session 1's critical-path question. Flat 4x4 VQ at k=256 was prototyped
and REJECTED by eye: Dirk's face disintegrates and ink outlines break into
4-pixel stair-steps. The 256-colour palettised frame is excellent, so the
palette was never the problem -- block VQ was.

Replaced it with a Cinepak-style hybrid: each 4x4 block is SKIP, one 4x4
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
rate-distortion. The RAW escape makes lam=0 pixel-exact (measured 0.00 dB loss),
so the quality knob spans lossless to heavily-compressed in one bitstream.

Per the user's decision, ships TWO quality profiles from that one codec, one
decoder and one bitstream -- only the rate knob differs:
  sasi  45 KB/s  lam=300  34.8 dB   stock 10MHz ACE/EXPERT
  scsi  75 KB/s  lam=100  35.9 dB   Super/XVI or CZ-6BS1

Three corrections to earlier numbers:

1. Session 1's "183 KB/s at 12fps" was a bad extrapolation. Halving the
   framerate does not halve the bitrate -- decimation roughly doubles the
   per-frame delta. Re-measured directly: 340 KB/s for session 1's own RLE,
   247 KB/s for changed-spans+deflate. The lossless floor is 319 MB.

2. A FOURTH false-good result, same family as the three in FINDINGS 4:
   k=1024 codebooks appeared to buy +2.4 dB free, because the rate model
   charged 1 byte for a 10-bit index. Charging the true cost reverses the
   verdict -- k=256 wins at every matched bitrate, and by 5 dB at the low end
   where the SASI profile lives. k=256 ships.

3. Stream inventory: the ~3-5MB clips are 1.2-1.7s, not ~60s, and some 60s
   streams are menus, not content. Any survey must classify before averaging.

Also cleared both candidate sources for the game-logic layer: the SNES project
is MIT and DirkSimple is zlib, so the arcade scene graph can be imported and
the two transcriptions diffed against each other.

Encoder is working end-to-end: extract.py -> vq/vq_hybrid/ratectl -> encode.py,
emitting a big-endian DLX1 container the 68000 can parse with plain moves.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 11:56:08 -07:00

6.7 KiB

Status & next-session handoff — end of session 2 (2026-08-23)

Decisions locked

decision value why
Target CPU 68000 @ 10MHz (stock) hardest honest constraint
Display mode 256 colors, 256x192 in 256x256 CRTC mode every mode is 1 word-access/pixel, so 256c is free vs 16c
Double buffer none — page 1 sacrificed enables movem.l 24px bursts; delta coding needs a RAM reference frame anyway
Codec hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion flat 4x4 VQ was measured and rejected — see FINDINGS 9-10
Quality modes two: sasi and scsi (USER DECISION, session 2) one codec, one decoder, one bitstream; only lam differs
Framerate 12 fps, explicit decimation source has zero duplicate frames; no free "twos" win
Emulator MAME 0.277 x68000 accurate enough that measured cycles mean something
SNES project reuse MIT — cleared data/events/ scene graph is reusable with attribution

The SASI/SCSI question is RESOLVED

Session 1 left "which machine do we target" open. The user's answer: ship both, as two quality profiles. This is now implemented rather than hypothetical — the bitrate ceiling is a build parameter in tools/encoder/ratectl.py:

profile target lam quality (00020 / 00146) machine
sasi 45 KB/s 300 34.8 / 28.3 dB stock 10MHz ACE/EXPERT
scsi 75 KB/s 100 35.9 / 29.0 dB Super/XVI, or CZ-6BS1 board

Codebooks are k=256 with 1-byte indices in both profiles. k=1024 was measured and rejected — see FINDINGS 14, it was a false-good result from a rate model that undercharged the index. Do not ship past lam~800; FINDINGS 15 has the cliff.

Because of the RAW escape mode, lam=0 is pixel-exact against the palettised frame (measured 0.00 dB loss). The profiles are two points on one continuous rate-distortion curve, not two codecs.


What session 2 settled

  1. The critical-path question is answered. "Does VQ soften Bluth's linework unacceptably?" — flat 4x4 k=256 VQ: yes, badly. Hybrid VQ with k=1024: no. Verified by eye, not just PSNR. See docs/FINDINGS.md 9-11.
  2. Session 1's 12fps bitrate was wrong (183 KB/s claimed, 340 KB/s measured). Halving the framerate does not halve the bitrate. FINDINGS 8. 2b. A fourth false-good result was produced and caught this session — k=1024 codebooks looked like a +2.4 dB free win because the rate model charged 1 byte for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
  3. The 256-colour palettised frame is the real quality ceiling and it looks excellent. Judge the codec against that, not against 1080p.
  4. Encoder exists and produces a real bitstream: tools/encoder/.

Encoder — working

python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
python3 tools/encoder/encode.py  /tmp/fr_00020 out.dlx --profile sasi --preview p.png
file role
extract.py .m2ts -> 256x192 PNGs, 12fps, spatial-only denoise
vq.py palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR
vq_hybrid.py the codec: 4 block modes + lagrangian mode decision
ratectl.py SASI/SCSI profiles, leaky-bucket rate control
encode.py CLI + DLX1 container writer

DLX1 container layout is documented in the encode.py docstring. All multi-byte fields are big-endian so the 68000 reads them with a plain move.

Known encoder gaps

  • Rate control is written but not yet wired into encode.py — the CLI uses a fixed lam from the profile. ratectl.encode_rate_controlled() exists and builds a lam-ladder per frame; it needs hooking up and validating.
  • Payload is not entropy-coded. Deflate on the payload should buy ~1.4x (measured on the lossless path, FINDINGS 8). LZ decode is cheap on a 68000.
  • Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
  • _paint is a Python per-block loop — fine for prototyping, slow for a full disc encode. Vectorise before the 224-stream run.

Working setup (unchanged from session 1, re-verified)

MAME ROMs~/mame/roms/x68000.zip. Must pass -bios ipl10.

mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3

Assemblertools/vasm/vasmm68k_mot -Fbin -o out.bin in.s

Blu-rayudisksctl loop-setup -r -f DRAGONS_LAIR.iso -> /media/reala-misaki/BDROM (still mounted as of end of session 2).

MAME Lua harnesstools/bench/*.lua, working. Three gotchas (retain the notifier subscription in a global; the stack register is SP not A7; autoboot_script fires at PC=0 before boot) are documented in FINDINGS.

Two shell traps, both hit again this session:

  • piping MAME (or any long job) through grep block-buffers — write to a file.
  • pkill -f <pattern> matches your own shell and kills it (exit 144). Use pkill -x or kill by PID.

STILL BLOCKED: disk throughput benchmark

Unchanged from session 1 — IOCS _B_READ returns -1 uniformly. Full diagnosis and the four untested hypotheses are in session 1's notes (git history of this file, commit 65112b9).

This now matters more than session 1 thought. Session 1 dismissed it because "VQ at 30 KB/s is correct whether SASI does 300 or 600 KB/s". But we now ship two profiles, and the profile bitrates (45 / 120 KB/s) are set against folklore bandwidth figures. A real measurement would let us set them honestly instead of conservatively. Next move is the untried SCSI path: -exp1 cz6bs1 -hard disk.chd.


Next steps, in priority order

  1. Wire rate control into encode.py and validate that the hard ceiling actually holds on an action scene (the whole point of choosing VQ).
  2. Entropy-code the payload (deflate) — ~1.4x for cheap 68000 decode cost.
  3. 68000 decoder skeleton: parse DLX1, expand codebooks to word-per-pixel, blit V1/V4/RAW/SKIP. Measure real cycles with the existing MAME Lua harness — this is the first time the harness gets used for its actual purpose.
  4. Full-disc survey — classify menu vs content first (FINDINGS 13), then measure bitrate across all 224 streams per profile.
  5. Resolve the framing question (FINDINGS 12: crop vs squash vs wide).
  6. Unblock the disk benchmark via the SCSI path, then re-set profile bitrates.
  7. Import the SNES project's data/events/ (MIT, cleared) as the scene graph. Cross-check against DirkSimple (zlib) which has the same data independently.
  8. ADPCM audio: MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in ratectl, not yet extracted or encoded.

Not yet started

  • Any 68000 player code
  • ADPCM audio extraction/encoding
  • Disk image packaging
  • Game logic (scene branching, input windows, death clips)