Answers session 1's critical-path question. Flat 4x4 VQ at k=256 was prototyped and REJECTED by eye: Dirk's face disintegrates and ink outlines break into 4-pixel stair-steps. The 256-colour palettised frame is excellent, so the palette was never the problem -- block VQ was. Replaced it with a Cinepak-style hybrid: each 4x4 block is SKIP, one 4x4 codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion. The RAW escape makes lam=0 pixel-exact (measured 0.00 dB loss), so the quality knob spans lossless to heavily-compressed in one bitstream. Per the user's decision, ships TWO quality profiles from that one codec, one decoder and one bitstream -- only the rate knob differs: sasi 45 KB/s lam=300 34.8 dB stock 10MHz ACE/EXPERT scsi 75 KB/s lam=100 35.9 dB Super/XVI or CZ-6BS1 Three corrections to earlier numbers: 1. Session 1's "183 KB/s at 12fps" was a bad extrapolation. Halving the framerate does not halve the bitrate -- decimation roughly doubles the per-frame delta. Re-measured directly: 340 KB/s for session 1's own RLE, 247 KB/s for changed-spans+deflate. The lossless floor is 319 MB. 2. A FOURTH false-good result, same family as the three in FINDINGS 4: k=1024 codebooks appeared to buy +2.4 dB free, because the rate model charged 1 byte for a 10-bit index. Charging the true cost reverses the verdict -- k=256 wins at every matched bitrate, and by 5 dB at the low end where the SASI profile lives. k=256 ships. 3. Stream inventory: the ~3-5MB clips are 1.2-1.7s, not ~60s, and some 60s streams are menus, not content. Any survey must classify before averaging. Also cleared both candidate sources for the game-logic layer: the SNES project is MIT and DirkSimple is zlib, so the arcade scene graph can be imported and the two transcriptions diffed against each other. Encoder is working end-to-end: extract.py -> vq/vq_hybrid/ratectl -> encode.py, emitting a big-endian DLX1 container the 68000 can parse with plain moves. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
17 KiB
Findings — session 1 (2026-08-23)
All numbers here are MEASURED unless marked ESTIMATE or FOLKLORE.
1. Source material
DRAGONS_LAIR.iso — 16 GB, UDF 2.x, decrypted (no AACS dir).
Loop-mounted read-only at /media/reala-misaki/BDROM via udisksctl loop-setup -r -f.
(7-Zip cannot read UDF 2.x; use the loop mount.)
- 224
.m2tsstreams, 1920x1080, MPEG-2, progressive, 23.976 fps - Size histogram: 47 <5MB, 138 5-50MB, 22 50-150MB, 14 150-400MB, 3 >400MB
- The 185 sub-50MB streams are the arcade branching scenes already split into individual clips — we get scene boundaries for free.
- Big streams are full-feature playthroughs: 00215 (1376s), 00216 (1151s), 00223 (566s)
- Typical scene clip ~60s (00203/00205/00199), some ~100s (00164/00212)
Gotcha: clip durations vary wildly. Always read format=duration and seek
relative to it. Seeking to a fixed offset silently yields 0 frames on short clips.
2. GVRAM layout [verified — see HARDWARE.md for source]
One 16-bit word per pixel position in EVERY color mode. Bit depth does not change VRAM bandwidth; it only subdivides the word.
addr = page_base + y*1024 + x*2 — adjacent pixels are 2 bytes apart in all modes.
Consequence: low bit depth buys no speed. 16-color mode is strictly worse than 256-color (same bus traffic, 1/16 the palette). Page-alias writes are hardware auto-masked, so 16-color needs no software read-modify-write — but it's still one word-access per pixel.
Chosen: 256 colors, 256x192 active area.
In 256-color mode P0=low byte, P1=high byte of each word. Sacrificing page 1 as a
double-buffer lets a move.l cover two pixel positions, enabling movem.l bursts
(12 regs = 48 bytes = 24 pixels). Identical blit cost to 65536-color mode but
half the on-disk data.
3. Content measurements (8 scenes sampled, 5s each at 40% into each clip)
| metric | mean | p90 |
|---|---|---|
| pixels changed / frame | 20.1% | 30.2% |
| blit cost | ~64k cycles | ~97k cycles |
| naive delta+RLE frame size | 15.5 KB | 19.6 KB |
Budget is 833,333 cycles/frame @ 12fps on a 10MHz 68000.
=> THE CPU IS NOT THE BOTTLENECK. I/O IS.
Blit uses under 8% of budget. The naive row-span+RLE codec achieves only 3.2:1, giving 365 KB/s / 470 MB at 24fps (~183 KB/s / 235 MB at 12fps).
Per-scene variance is extreme: static dialogue ~30 KB/s, action ~700 KB/s. Any codec needs a hard bitrate ceiling, not just a good average.
"Shot on twos" — ASSUMPTION FAILED
Dedupe found zero duplicate frames across all 8 scenes (uniq=120/120,
24.0 fps effective). This Blu-ray is a restoration where every frame is unique.
We do NOT get halved data for free. Decimation to 12fps must be explicit.
A weak alternation signature does exist (even-index pairs 40.7% vs odd 27.5%, ratio 1.5x, with occasional true-duplicate pairs at 0.03-0.19%), but it is irregular — Bluth mixed ones and twos; action is animated on ones.
4. MEASUREMENT TRAPS — read before trusting any pipeline number
Three separate false results were produced and caught this session. All three looked plausible. Guard against them:
-
Per-frame Floyd-Steinberg dithering destroys temporal coherence. Error diffusion is chaotic: a +/-1 input change cascades across the row and produces a completely different index pattern. First run reported 31.5% pixels changed with near-zero variance (median 31.6, p90 32.3, max 32.7) while source mean-abs-diff was 0.09 — i.e. visually identical frames. That flat variance is the tell: real animation has scene-dependent variance; noise does not. Use no dithering (cel art is flat) or ordered/Bayer (spatially fixed, temporally stable).
-
Temporal denoise smears motion.
hqdn3d=4:3:6:4— the6:4are temporal params. It flattened real motion, which then measured as "no motion" and produced an absurd 0.8 fps / 4 MB result. Use spatial-only:hqdn3d=4:3:0:0. -
Exact-match dedupe fails on a noisy source. MPEG-2 grain means near-duplicate frames differ by +/-1 and are never bit-exact. Use a threshold on "% pixels differing by more than N levels", and pick the threshold from the observed distribution, not a guess. A 2% threshold ate genuine animation when mean consecutive change was only 0.9%.
Sanity rule: if a result has suspiciously low variance, or is suspiciously good, it is probably an artifact of the measurement, not a property of the content.
Scripts kept in tools/analysis/ — 01 and 02 are marked BROKEN deliberately as
regression references; 03 and 04 are the correct ones.
5. Storage interface — the SASI/SCSI split
[Yasuma, X68030 internal SCSI controller]
- Interface: SCSI-1, 50-pin, 5 MB/s bus spec
- Controller: Fujitsu MB89352 SPC
- Transfer mode: DMA (via HD63450 DMAC)
- Bus: X68000 original bus, 16-bit @ 10MHz
Even on the X68030, SCSI runs at 10MHz 16-bit DMA. Storage bandwidth does NOT scale with CPU — the controller sits on the original bus. HD63450's 12.5MHz official ceiling is why the X68030 runs at 25MHz. An "HSCSI" TSR forces PIO/FIFO transfer instead of DMA but was marginal even at 25MHz.
Because it's DMA, streaming costs essentially no CPU — this stacks with the 8% blit utilisation. The 68000 really is nearly idle.
Model split — IMPORTANT
The 10MHz models (original X68000, ACE, PRO, EXPERT) use SASI, not SCSI.
Built-in SCSI starts at the X68000 Super (1990) and continues through XVI,
Compact, X68030. SCSI on earlier machines needs the Sharp CZ-6BS1 board
in an I/O slot (MAME models this: -exp1 cz6bs1).
| target | bandwidth | naive codec (365 KB/s) | VQ codec (~30 KB/s) |
|---|---|---|---|
| SASI (stock ACE/EXPERT) | ~300-500 KB/s FOLKLORE | infeasible | comfortable |
| SCSI (Super+, or CZ-6BS1) | ~1 MB/s FOLKLORE | tight but viable | trivial |
Derived bounds (ESTIMATE): 16-bit @10MHz with 4-clock bus cycle = 5 MB/s absolute ceiling; HD63450 single-address DMA ~8 clocks/word => ~2.5 MB/s practical ceiling, before SCSI-1 async handshake and drive latency.
No measured benchmark was obtained — see STATUS.md. The ~300-500 KB/s and ~1 MB/s figures are folklore-grade; I could not find a primary measurement.
6. Codec decision: vector quantization (Cinepak-style)
Given ~8x CPU headroom and an I/O ceiling, spend CPU to buy bandwidth.
- Split frame into 4x4 blocks, encode each as a 1-byte index into a per-scene codebook
- Decode = 16-byte copy from a lookup table: nearly free
- A full frame = 256*192/16 = 3,072 bytes — a hard 16:1 floor before delta
- Add block-level delta on top; action scenes ~2-3 KB/frame
- => roughly 30 KB/s, ~40 MB total, with a deterministic bitrate ceiling
Divergence from the SNES project (below): use a per-scene codebook with delta updates, not a per-frame rebuild. We trade adaptivity for bandwidth because we have 2MB RAM to keep a codebook resident and CPU to spare.
Risk not yet evaluated: 4x4 VQ with a 256-entry codebook will visibly soften detail. Bluth's fine ink linework is what suffers. Prototype and eyeball before committing.
7. Comparison: astrobleem/SNES-SuperDragonsLairArcade
Reached the same core architecture independently — "512 tiles per frame" is vector quantization (8x8 codebook + tilemap). Good validation.
But: the SNES PPU has no bitmap mode, so tiles are forced on them by display hardware. The X68000 has a real linear framebuffer, so VQ is a compression choice we can tune or drop per-scene.
MSU-1 is a bandwidth cheat we don't have. It's a modern flash-cart coprocessor giving memory-mapped streaming the real SNES never had. Their budget: 512 tiles x 32 bytes (4bpp 8x8) + tilemap ~= 18 KB/frame => ~430 KB/s at 23.976fps. That's higher than the 365 KB/s we'd reject on SASI. (ESTIMATE: my arithmetic on their stated tile budget, not a measured figure.)
Where we're ahead: 256 simultaneous colors from a 65536 palette vs their 4bpp sub-palettes needing a tile-aware palette optimizer plus a spatial smoothing pass to hide 8x8 palette seams. That problem doesn't exist for us. Plus 68000@10MHz vs 65816@3.58MHz, and 2MB vs 128KB.
Most valuable thing in that repo is NOT the codec — it's data/events/:
516 chapter definitions across 29 scenes as XML, plus
data/chapter_event_inventory.md. That's the arcade scene graph and input-timing
structure, entirely hardware-independent — the whole game-logic layer we'd
otherwise reverse-engineer from the arcade ROM.
TODO: check their license before planning to reuse it. Their 516 chapters are finer-grained than our 224 Blu-ray streams, so mapping their event table onto our footage means subdividing streams by timecode.
Caveat: all of the above is from README/repo-tree summaries, not their source.
Findings — session 2 (2026-08-23)
8. CORRECTION to session 1: halving the framerate does NOT halve the bitrate
Session 1 measured 365 KB/s for naive delta+RLE at 24 fps and wrote "(~183 KB/s at 12fps)". That extrapolation is wrong. Decimating to 12 fps roughly doubles the per-frame delta, so the rate stays nearly flat.
Re-measured directly on 12 fps decimated frames (4 scenes, 66 frames):
| codec (all LOSSLESS w.r.t. the 256-colour frame) | B/frame | KB/s @12 | 22 min | ratio |
|---|---|---|---|---|
| raw 8bpp 256x192 | 49152 | 576 | 743 MB | 1.0:1 |
| session 1 row-span + RLE | 29055 | 340 | 439 MB | 1.7:1 |
| XOR vs prev + deflate | 30196 | 354 | 456 MB | 1.6:1 |
| changed-spans + deflate | 21110 | 247 | 319 MB | 2.3:1 |
| changed-spans + LZMA | 18759 | 220 | 283 MB | 2.6:1 |
Session 1's own RLE re-measured at 12 fps gives 340 KB/s, not 183. Any plan that assumed 183 KB/s was based on a bad number.
Deflate-class entropy coding on top of the span payload is worth 1.4x over hand-rolled RLE, and LZ decode is cheap on a 68000 (byte copies), so the lossless floor is ~247 KB/s / 319 MB. That is infeasible on SASI and tight but real on SCSI.
9. Flat 4x4 VQ at k=256 is NOT acceptable — confirmed by eye
The risk flagged in 6 is real. At k=256, 4x4:
| scene | palette-only PSNR | after VQ | VQ loss |
|---|---|---|---|
| 00010 | 38.35 | 29.68 | 8.67 dB |
| 00020 | 39.90 | 32.67 | 7.22 dB |
| 00146 | 35.25 | 29.35 | 5.89 dB |
| 00181 | 41.92 | 32.87 | 9.05 dB |
Visually: Dirk's face disintegrates, teeth and eyes turn to mush, ink outlines break into 4-pixel stair-steps, colour bleeds across block boundaries.
Left: 1080p source. Middle: 256-colour palettised 256x192 — the quality ceiling,
and it is excellent. Right: flat 4x4 VQ at k=256. This is the result that killed
the flat-VQ architecture.
Crucially, the 256-colour palettised frame itself looks excellent. Flat cel art with a per-scene median-cut palette and no dithering is near-transparent (35-42 dB). So the palette is not the problem and 256 colours is not the problem — block VQ is. The quality ceiling we should hold ourselves to is the palettised frame, not the 1080p source.
10. Hybrid VQ (Cinepak V1/V4 + SKIP) — this is the codec
Per 4x4 block, choose by rate-distortion: SKIP (reuse previous frame),
V1 (one 4x4 codeword, 1 byte), or V4 (four 2x2 codewords, 4 bytes),
with a 2-bit-per-block mode header. lam is the lagrangian rate knob.
Measured, k1=k4=256, 4 scenes (mean of the per-scene table in the session log):
| lam | PSNR | loss vs palette | SKIP% | V1% | V4% | B/frame | KB/s @12 |
|---|---|---|---|---|---|---|---|
| 0 (max quality) | 33.9 | 4.9 | 30.8 | 18.5 | 50.8 | 7574 | 88.8 |
| 200 | 31.9 | 5.9 | 44.0 | 37.6 | 18.4 | 4183 | 49.0 |
| 1000 | 31.6 | 7.3 | 47.4 | 47.7 | 4.9 | 2841 | 33.3 |
| 5000 | 25.5 | 13.3 | 55.6 | 44.4 | 0.0 | 2134 | 25.0 |
At a matched ~30 KB/s the hybrid beats flat 4x4 VQ by ~1 dB, and unlike flat VQ it keeps scaling: at 89 KB/s it reaches within 4.9 dB of the palette ceiling, which flat VQ cannot reach at any bitrate.
Note V4% collapses to 0 at lam=5000 — that is the knob doing exactly what it should: under a hard ceiling, detail blocks are the first thing sacrificed.
11. Codebook size sweep (flat 4x4, for reference)
| block | k | PSNR | loss | key B | changed% | KB/s @12 | codebook RAM |
|---|---|---|---|---|---|---|---|
| 4x4 | 256 | 30.46 | 8.39 | 3072 | 52.7 | 28.5 | 8K |
| 4x4 | 1024 | 32.89 | 5.96 | 3840 | 56.6 | 35.6 | 32K |
+2.4 dB for 24K more RAM and 7 KB/s. With 2 MB of RAM, a 1024-entry codebook is cheap and clearly worth it. (RAM figure is the word-expanded form the blitter wants: k * 16 px * 2 bytes.)
12. Source framing — OPEN
The Blu-ray is full-frame 1920x1080 16:9 with no pillarboxing. The arcade
original is 4:3. The extractor currently centre-crops 1440x1080, which is the
arcade-faithful choice but discards image the 2006 remaster added. Options are
crop (default), squash, wide in tools/encoder/extract.py.
Not yet decided; needs an eyeball comparison against arcade reference.
13. Stream inventory correction
Session 1 said "typical scene clip ~60s". Sampled directly: the ~3-5 MB streams are 1.2-1.7 s clips — these are the individual arcade death/action moments, which is exactly the granularity the game logic needs. Some 60 s streams (e.g. 00203) are menu screens, not content. Any survey must classify menu vs content before averaging, or the bitrate numbers are diluted by static menus.
14. A FOURTH false-good result — and the correction
Add this to the 4 list. The mechanism was new but the shape was identical.
The false result: flat and hybrid VQ both showed +2.4 dB for k=1024 over k=256 at an apparently similar bitrate, which made a 1024-entry codebook look like an obvious win. The k=1024 quality ladder rendered from that run looked great at "45 KB/s".
The bug: the rate-distortion model in vq_hybrid.encode() charged 1 byte
per codebook index unconditionally. A 1024-entry codebook needs a 10-bit index,
stored as 2 bytes. So every k=1024 measurement understated the V1 and V4
payload by exactly 2x, and the lagrangian mode decision was choosing V4 on the
belief that four codewords cost 4 bytes when they cost 8.
After charging the true index cost (idx_bytes is now explicit and defaults
from the codebook size), matched-bitrate comparison on scene 00020:
| KB/s | k=256 (1-byte idx) | k=1024 (2-byte idx) |
|---|---|---|
| ~32-42 | 33.87 dB @ 32.5 | 28.91 dB @ 42.3 |
| ~44-52 | 34.80 dB @ 44.1 | 35.13 dB @ 52.5 |
| ~72-86 | 35.87 dB @ 72.2 | 36.51 dB @ 86.0 |
k=1024 buys +0.3 to +0.6 dB for +19% bitrate — a wash at best — and at the low end where the SASI profile lives it is 5 dB worse, because the 2-byte index floor dominates once V4 is priced out.
k=256 with 1-byte indices is the shipping choice. It is also the better
decoder: a plain move.b index with no alignment case, and an 8 KB codebook
instead of 32 KB.
The general lesson, again: the comparison was not wrong about VQ, it was wrong about cost. When a knob looks like a free win, check that the rate model is charging for it. Same failure family as 4.1-4.3: a plausible number produced by a pipeline that was not measuring what it claimed to measure.
15. Rate-distortion curve of the shipping codec (k=256, corrected)
Scene 00020 (Dirk screaming, close-up face — the hardest case for linework), and 00146. Includes the 2-bit-per-block mode header. No entropy coding yet.
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | SKIP | V1 | V4 | RAW |
|---|---|---|---|---|---|---|---|---|
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~37% | ~24% | ~13% | ~26% |
| 100 | 35.87 | 72.2 | 29.04 | 72.5 | ~41% | ~34% | ~21% | ~4% |
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | ~44% | ~42% | ~14% | 0% |
| 800 | 33.87 | 32.5 | 27.77 | 36.1 | ~46% | ~48% | ~5% | 0% |
| 2000 | 27.57 | 25.5 | 24.88 | 30.2 | ~50% | ~49% | ~1% | 0% |
Palette ceilings: 00020 = 39.90 dB, 00146 = 35.25 dB.
The shipping codec across the rate knob. Top: source, palette ceiling, lam=25.
Bottom: lam=100 (scsi profile), lam=300 (sasi profile), lam=800.
Both shipping profiles hold Bluth's linework; the failure only starts past lam=800.
Two things to read off this table:
- The cliff is between lam=800 and lam=2000. That is where V4 is priced out entirely and detail blocks have nowhere to go. Do not ship past lam~800.
- RAW is doing real work at high bitrate (26% of blocks at lam=25) and vanishes by lam=300. It is what makes the top of the curve reach the palette ceiling, and it costs the decoder nothing — RAW is the cheapest mode to blit.
16. Licences cleared for the game-logic layer
Both checked this session:
- astrobleem/SNES-SuperDragonsLairArcade — MIT, "Copyright (c) 2026 Chad
Doebelin".
data/events/holds 516 XML chapter definitions with timing and event data. Reusable with attribution. - icculus/DirkSimple — zlib. Independent from-scratch reimplementation of
the game logic in Lua, scene/timing tables in
game.lua. Also permissive.
Having two independent permissively-licensed transcriptions of the arcade scene graph is better than one: they can be diffed against each other to catch transcription errors before any of it is committed to 68000 tables.