The codec was designed when bytes were scarce, so every decision in it trades
cycles to save bytes. That is now backwards: sasi spends 110 KB/s of a 488 KB/s
pipe while missing 31% of frames on CPU.
The cheapest thing a 68000 can be handed is the most expensive thing to store.
Measured, per pixel: row-linear copy from word-expanded memory 9.08 cycles,
block-order 12.98, V1 codebook 18.74, RAW byte literals 25.03. So the 1024-byte
stride costs 43% and unpacking bytes to words costs more than the write itself.
Pricing one new mode -- a per-row span of word-expanded literals movem.l'd
straight from the stream buffer -- against the UNCHANGED mode maps:
sasi median 74.4% -> 43.0%, worst 136.2% -> 106.2%, misses 37 -> 8/120,
101.7 -> 453.2 KB/s
scsi median 94.9% -> 69.4%, misses 51 -> 18/120, 272 -> 479.7 KB/s
scsi gains less precisely because it has less idle bandwidth left to trade.
Two consequences worth flagging. A word-expanded literal block derives to ~240
cycles, cheaper than V1's measured 299.9 and pixel-exact -- so every codebook
mode is CPU-dominated by a literal, and the codebook is a byte optimisation
that now costs cycles. And 28.5's "a scene cut cannot fit at 12fps" reopens:
CPU needs >=19% of the frame as spans, the bus allows <=39%, and that interval
is not empty.
DERIVED, NOT MEASURED, and labelled as such everywhere. The 9.08 cycles/pixel
is real but was measured at full row width with 12-register bursts, so short
spans are flattered. Measuring one span on the 68000 is now step 0 of the next
session, ahead of the cost-aware mode decision, because it changes the mode set
that decision optimises over.
FINDINGS 29. tools/analysis/12_span_tradeoff.py.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Dragon's Lair — Sharp X68000 port
Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SASI/SCSI).
This is fundamentally a video codec problem, not a game-logic problem: the game logic is a scene table with branching input windows; the difficulty is pushing ~22 minutes of Don Bluth animation through a 10MHz 68000.
Green-light check: ./tools/bench/check.sh (~3 min, needs the Blu-ray
mounted) re-runs both display regression tests, the rate-control drift test, the
display-path coherency counterexample and a 120-frame 68000 decode, then prints
ALL GREEN.
Read first
docs/FINDINGS.md— measured hardware facts, content statistics, codec decision, and a section on measurement traps that produced three separate false results. Read §4 before trusting any pipeline number.docs/STATUS.md— current state, working setup, blockers, next steps. Start here. It also lists what has been explicitly abandoned, so old ideas do not get re-proposed.docs/BENCHMARK.md— how to measure the storage subsystem, and why a bandwidth figure out of MAME would be meaningless.docs/HARDWARE.md— X68000 GVRAM/CRTC reference.
Layout
docs/ findings, status, hardware reference
tools/analysis/ measurement scripts, numbered in the order they were written
(01/02 marked BROKEN deliberately, kept as regression refs).
Run from the repo root — they import from tools/encoder/.
07 finds the hottest sustained window in a stream; 08 renders
source | decoded | block-mode map as .webm; 09 is the
rate-control drift gate (FINDINGS 26/27) and is part of
check.sh -- it exits non-zero if the encoder ever again
reports a reconstruction no decoder would produce.
10 is a COUNTEREXAMPLE, and exits non-zero by design: it
demonstrates that the two-display-path plan of FINDINGS
24.5/25.6 corrupts 70 of 120 frames (FINDINGS 28.1).
11 scores a container against the MEASURED per-mode block
costs without needing MAME.
tools/bench/ MAME Lua injection harness + 68000 benchmark sources.
`check.sh` re-runs both display regression tests (~40 s).
`blit.s`/`blit.lua` time the full-frame GVRAM blit on the
68000 itself (FINDINGS 24) — not part of check.sh, because
wall timings would make the green-light check host-sensitive.
`crtc_mode.lua` is the single source of truth for CRTC R00-R08
and R20 — do not write CRTC values anywhere else.
`prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and
verify `src/player/decode.s`; the verify pass is in check.sh.
tools/vasm/ vasm m68k assembler (built from source)
tools/encoder/ hybrid VQ encoder + DLX1 container writer (working).
dlx.py is the reference DECODER -- ground truth for the 68000.
src/player/ decode.s: the 68000 DLX1 decoder. Pixel-exact, and 31% of
frames over the 12fps CPU budget. See FINDINGS 28.
assets/ extracted frames/audio (gitignored)
Encoder
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile sasi --preview p.png
Two quality profiles ship from one codec and one decoder — sasi (110 KB/s) and
scsi (280 KB/s) are two points on the same rate-distortion curve. Both are
ceilings: lam is bisected per frame under a leaky bucket, so the profile's
lam is a quality floor rather than a setting (--fixed-lam opts out). The codec is
a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four
2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion.
The RAW escape means lam=0 is pixel-exact against the palettised frame, so the
quality knob spans lossless to heavily-compressed without changing the bitstream.
Profiles are derived from a bandwidth figure, not chosen by eye:
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
On reading
docs/FINDINGS.md: it is append-only and several later sections overturn earlier ones. Superseded sections carry a blockquote at the top pointing to the correction — heed those, especially 18 (reversed by 21).
Source media (DRAGONS_LAIR.iso) and ROMs are gitignored — supply your own.
Not every large stream is game footage. 00216 is the feature with a
burned-in commentary picture-in-picture and 00215 is the commentary itself —
the two largest files on the disc. The clean 9.4-minute animation is 00223.
See FINDINGS 25.1 before running any size-ranked survey.