prosolis e1aa26bb57 The 68000 decoder draws pixel-exact frames, and does not fit
src/player/decode.s parses DLX1 and decodes straight into GVRAM. Verified
pixel-exact over a 120-frame sequential run of the worst sustained window on
the disc -- all four block modes, full temporal recursion, so the last frame
is only right if all 120 were. In check.sh.

It costs a mean of 81.7% of a 12fps frame budget, and 31% of frames exceed
100% (42% at scsi). CPU is now the binding constraint. FINDINGS 28.

Three things that were believed and are not true:

- The dual-display-path plan of FINDINGS 24.5/25.6 is incoherent. The compose
  path needs a RAM copy of the previous reconstruction; the direct path's
  selling point is that it keeps none. Mixing them shows stale pixels on 70 of
  120 frames, worst frame 18.8% of the screen. Every coherent repair is dearer
  than not mixing, and 24.5's two figures were both copies with no decode in
  either, so there was never a crossover to find. One path ships, and the 96KB
  reference frame is gone. tools/analysis/10_pathmix_drift.py keeps the
  counterexample runnable; check.sh asserts it still reproduces.

- The four block modes do not cost the same. V1 300, V4 448, RAW 400 cycles
  against the old model's flat 207.8. V4 is 25% of blocks and 50% of the
  cycles, and the mode decision charges it bytes it does not charge cycles for.
  tools/analysis/11_cpu_budget.py reproduces all four frames timed on the
  68000 to within 1 point. Hand-derived timings agree to 0.5% on V1.

- The container is big-endian but not aligned. Variable-length records laid end
  to end put frame 1's length field at an odd address, and move.l (a0)+ there
  is an address error: frame 0 decoded perfectly and then vectored into the
  IPL for 59 emulated seconds looking like a hang. Found by dumping PC, not by
  reading the source.

Also: an all-V1 frame, the cheapest possible full redraw, is 110.5% of budget.
No mode assignment fits a scene cut at 12fps. That one needs a decision, not a
measurement.

Next: charge cycles in the mode decision and bisect against 833,333 per frame,
the way session 6 bisects lam against bytes -- but with no bucket, because a
late frame cannot be banked.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 15:04:38 -07:00

Dragon's Lair — Sharp X68000 port

Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SASI/SCSI).

This is fundamentally a video codec problem, not a game-logic problem: the game logic is a scene table with branching input windows; the difficulty is pushing ~22 minutes of Don Bluth animation through a 10MHz 68000.

Green-light check: ./tools/bench/check.sh (~3 min, needs the Blu-ray mounted) re-runs both display regression tests, the rate-control drift test, the display-path coherency counterexample and a 120-frame 68000 decode, then prints ALL GREEN.

Read first

  • docs/FINDINGS.md — measured hardware facts, content statistics, codec decision, and a section on measurement traps that produced three separate false results. Read §4 before trusting any pipeline number.
  • docs/STATUS.md — current state, working setup, blockers, next steps. Start here. It also lists what has been explicitly abandoned, so old ideas do not get re-proposed.
  • docs/BENCHMARK.md — how to measure the storage subsystem, and why a bandwidth figure out of MAME would be meaningless.
  • docs/HARDWARE.md — X68000 GVRAM/CRTC reference.

Layout

docs/            findings, status, hardware reference
tools/analysis/  measurement scripts, numbered in the order they were written
                 (01/02 marked BROKEN deliberately, kept as regression refs).
                 Run from the repo root — they import from tools/encoder/.
                 07 finds the hottest sustained window in a stream; 08 renders
                 source | decoded | block-mode map as .webm; 09 is the
                 rate-control drift gate (FINDINGS 26/27) and is part of
                 check.sh -- it exits non-zero if the encoder ever again
                 reports a reconstruction no decoder would produce.
                 10 is a COUNTEREXAMPLE, and exits non-zero by design: it
                 demonstrates that the two-display-path plan of FINDINGS
                 24.5/25.6 corrupts 70 of 120 frames (FINDINGS 28.1).
                 11 scores a container against the MEASURED per-mode block
                 costs without needing MAME.
tools/bench/     MAME Lua injection harness + 68000 benchmark sources.
                 `check.sh` re-runs both display regression tests (~40 s).
                 `blit.s`/`blit.lua` time the full-frame GVRAM blit on the
                 68000 itself (FINDINGS 24) — not part of check.sh, because
                 wall timings would make the green-light check host-sensitive.
                 `crtc_mode.lua` is the single source of truth for CRTC R00-R08
                 and R20 — do not write CRTC values anywhere else.
                 `prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and
                 verify `src/player/decode.s`; the verify pass is in check.sh.
tools/vasm/      vasm m68k assembler (built from source)
tools/encoder/   hybrid VQ encoder + DLX1 container writer (working).
                 dlx.py is the reference DECODER -- ground truth for the 68000.
src/player/      decode.s: the 68000 DLX1 decoder. Pixel-exact, and 31% of
                 frames over the 12fps CPU budget. See FINDINGS 28.
assets/          extracted frames/audio (gitignored)

Encoder

python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
python3 tools/encoder/encode.py  /tmp/fr out.dlx --profile sasi --preview p.png

Two quality profiles ship from one codec and one decoder — sasi (110 KB/s) and scsi (280 KB/s) are two points on the same rate-distortion curve. Both are ceilings: lam is bisected per frame under a leaky bucket, so the profile's lam is a quality floor rather than a setting (--fixed-lam opts out). The codec is a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion.

The RAW escape means lam=0 is pixel-exact against the palettised frame, so the quality knob spans lossless to heavily-compressed without changing the bitstream.

Profiles are derived from a bandwidth figure, not chosen by eye:

python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi

On reading docs/FINDINGS.md: it is append-only and several later sections overturn earlier ones. Superseded sections carry a blockquote at the top pointing to the correction — heed those, especially 18 (reversed by 21).

Source media (DRAGONS_LAIR.iso) and ROMs are gitignored — supply your own.

Not every large stream is game footage. 00216 is the feature with a burned-in commentary picture-in-picture and 00215 is the commentary itself — the two largest files on the disc. The clean 9.4-minute animation is 00223. See FINDINGS 25.1 before running any size-ranked survey.

S
Description
No description provided
Readme
25 MiB
Languages
Python 53.1%
Assembly 20.4%
Lua 15.3%
Shell 8.8%
C 2.3%
Other 0.1%