blit.s gains v7 -- v6's 24-pixel movem chain plus a second chain whose unit is
one `move.l (a0)+,(a2)+`. Measured over 13 span lengths by span.sh, every config
pixel-exact:
cycles = 66.0 per span + 9.143 per COARSE pixel + 9.978 per FINE pixel
fitting all 13 to within 0.2%. v5 and v6 re-measure to FINDINGS 30 exactly, so
the harness has not drifted underneath the new variant.
Rescored against the same scsi window and the same additive model, v7 takes
84/120 frames over budget to 18/120 -- exactly what FINDINGS 39.4 derived, and
that agreement is two cancelling errors: the derivation's 2-register movem tail
is 29% too dear per pixel, and its "nothing per span" for the second chain entry
is 22.3 clocks too cheap. The plain post-incrementing move.l is the right tail
instruction, and it makes the padding quantum 2 pixels, which a run of 4x4
blocks pads to exactly zero.
The DMAC stays dropped on a measurement now rather than an argument: v7 takes
back 37 of the 43 frames the array chain would, with no reserved channel and no
timing neither emulator here can verify. Break-even against all-V1 moves from
L=4 blocks to L=2.
The fine displacement is carried mid-stream rather than in the span record, so
the decoder holds nothing across the copy and keeps all 12 payload registers --
which is the whole reason the coarse unit is 24 pixels.
span.sh is now -seconds_to_run 200 (30 s wall, 36 configs) and takes its
expected snapshot count from the generated metadata instead of a literal 23.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
144 lines
8.2 KiB
Markdown
144 lines
8.2 KiB
Markdown
# Dragon's Lair — Sharp X68000 port
|
|
|
|
Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SCSI).
|
|
|
|
This is fundamentally a **video codec problem**, not a game-logic problem: the
|
|
game logic is a scene table with branching input windows; the difficulty is
|
|
pushing ~22 minutes of Don Bluth animation through a 10MHz 68000.
|
|
|
|
**And the binding resource is the 68000's local BUS, not its clock.** The
|
|
decoder occupies 86.7% of it once instruction prefetch is counted, and 52 of the
|
|
53 frames that miss the 12fps budget miss it on the bus, not the CPU
|
|
(FINDINGS 38). Read that before optimising anything for cycles.
|
|
|
|
The largest measured win on the table is the **literal span with a fine tail**
|
|
(`blit.s` v7): it takes the worst `scsi` window from 84/120 frames over budget
|
|
to 18/120, and `src/player/decode.s` does not implement it yet (FINDINGS 40).
|
|
|
|
**Green-light check:** `./tools/bench/check.sh` (~3 min, needs the Blu-ray
|
|
mounted) re-runs both display regression tests, the rate-control drift test, the
|
|
display-path coherency counterexample and a 120-frame 68000 decode, then prints
|
|
`ALL GREEN`.
|
|
|
|
## Read first
|
|
- **`docs/FINDINGS.md`** — measured hardware facts, content statistics, codec
|
|
decision, and a section on measurement traps that produced three separate
|
|
false results. Read §4 before trusting any pipeline number.
|
|
- **`docs/STATUS.md`** — current state, working setup, blockers, next steps.
|
|
**Start here.** It also lists what has been explicitly abandoned, so old ideas
|
|
do not get re-proposed.
|
|
- **`docs/BENCHMARK.md`** — how to measure the storage subsystem, and why a
|
|
bandwidth figure out of MAME would be meaningless.
|
|
- **`docs/HARDWARE.md`** — X68000 GVRAM/CRTC reference.
|
|
|
|
## Layout
|
|
```
|
|
docs/ findings, status, hardware reference
|
|
tools/analysis/ measurement scripts, numbered in the order they were written
|
|
(01/02 marked BROKEN deliberately, kept as regression refs).
|
|
Run from the repo root — they import from tools/encoder/.
|
|
07 finds the hottest sustained window in a stream; 08 renders
|
|
source | decoded | block-mode map as .webm; 09 is the
|
|
rate-control drift gate (FINDINGS 26/27) and is part of
|
|
check.sh -- it exits non-zero if the encoder ever again
|
|
reports a reconstruction no decoder would produce.
|
|
10 is a COUNTEREXAMPLE, and exits non-zero by design: it
|
|
demonstrates that the two-display-path plan of FINDINGS
|
|
24.5/25.6 corrupts 70 of 120 frames (FINDINGS 28.1).
|
|
11 scores a container against the MEASURED per-mode block
|
|
costs without needing MAME; 12 prices the literal-span mode of
|
|
FINDINGS 30 against those same mode maps, and prints whether a
|
|
scene cut still fits at 12fps; 13 measures what fitting the
|
|
CPU budget costs in dB (FINDINGS 31) and caches H.build so the
|
|
search loop is seconds, not minutes.
|
|
14 prices the HD63450 array-chain against the v6 and v7
|
|
spans (FINDINGS 39/40) and prints the sensitivity that decides
|
|
it -- v7 is measured, and takes 37 of the 43 frames the DMAC
|
|
would, so the DMAC stays dropped;
|
|
15 measures how much of the 68000's LOCAL bus the decoder
|
|
occupies (FINDINGS 38) and exits non-zero if its derived
|
|
model stops matching the harness's measurement.
|
|
buscost.py is the shared bus-cycle table both import.
|
|
tools/bench/ MAME Lua injection harness + 68000 benchmark sources.
|
|
`check.sh` re-runs both display regression tests (~40 s).
|
|
`blit.s`/`blit.lua` time the full-frame GVRAM blit on the
|
|
68000 itself (FINDINGS 24) — not part of check.sh, because
|
|
wall timings would make the green-light check host-sensitive.
|
|
`span.sh` (prep_spans.py + span.lua + blit.s v5/v6/v7)
|
|
measures the literal-span mode the same way (FINDINGS 30 and
|
|
40, ~30 s); it also asserts that every one of its 36 timing
|
|
configs drew a pixel-exact frame, the count taken from the
|
|
generated metadata so a new config cannot weaken the gate.
|
|
v7 is v6 with a second, 2-pixel chain for the span tail:
|
|
66.0 cycles/span + 9.143 per coarse pixel + 9.978 per fine
|
|
pixel, MEASURED, which is the win FINDINGS 39.4 predicted.
|
|
`crtc_mode.lua` is the single source of truth for CRTC R00-R08
|
|
and R20 — do not write CRTC values anywhere else.
|
|
`prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and
|
|
verify `src/player/decode.s`; the verify pass is in check.sh.
|
|
tools/bench/c68k/ headless px68k C68K harness -- a SECOND emulator for every
|
|
68000 cycle figure (FINDINGS 37). Links only px68k's CPU core:
|
|
no SDL, no ROMs, no emulated machine. `make PX68K=~/src/px68k`
|
|
then `run.sh`; `verify_c68k.py` checks the decode is
|
|
pixel-exact, which is what licenses the cycle numbers. It also
|
|
counts BUS cycles, which MAME cannot report.
|
|
The Makefile's -no-pie and the harness's MAP_32BIT arena are
|
|
load-bearing: C68K truncates host pointers to 32 bits.
|
|
tools/vasm/ vasm m68k assembler (built from source)
|
|
tools/encoder/ hybrid VQ encoder + DLX2 container writer (working).
|
|
DLX2 4-byte-aligns every frame record: an odd `move.l` is an
|
|
ADDRESS ERROR on a 68000, not a slow read (FINDINGS 28.3).
|
|
dlx.py is the reference DECODER -- ground truth for the 68000.
|
|
src/player/ decode.s: the 68000 DLX decoder. Pixel-exact; 1 frame of 120
|
|
over the 12fps CPU budget once the mode decision prices
|
|
cycles. See FINDINGS 28 and 31.
|
|
assets/ extracted frames/audio (gitignored)
|
|
```
|
|
|
|
## Encoder
|
|
|
|
```
|
|
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
|
|
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile scsi --preview p.png
|
|
```
|
|
|
|
**One profile: `scsi`, 280 KB/s.** The 110 KB/s `sasi` profile was dropped in
|
|
session 9 on capacity, not bandwidth — a SASI volume is limited to 40 MB, and
|
|
the game's 22.8 minutes of footage is 146 MiB even at that rate (FINDINGS 32).
|
|
The rate point may return under another name once the delivery medium is
|
|
settled, because a 1x CD-ROM sustains ~150 KB/s and CD-ROM is the only period
|
|
medium with the capacity.
|
|
|
|
The profile bitrate is a **ceiling**: lam is bisected per frame under a leaky
|
|
bucket, so the profile's `lam` is a quality floor rather than a setting
|
|
(`--fixed-lam` opts out).
|
|
|
|
There are **two** ceilings, on two different axes. The second is the 68000's
|
|
decode budget: `mu` is bisected per frame against 833,333 cycles so the frame
|
|
also *decodes* in time, which takes the worst sustained window from 37 frames
|
|
over budget to 1 for 0.62 dB at `scsi` (FINDINGS 31). It is on by default; `--no-cpu-fit`
|
|
restores session 7 behaviour. Unlike bytes, cycles have no bucket — there is no
|
|
double buffer to decode ahead into, so it is a hard per-frame ceiling. The codec is
|
|
a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four
|
|
2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion.
|
|
|
|
The RAW escape means `lam=0` is pixel-exact against the palettised frame, so the
|
|
quality knob spans lossless to heavily-compressed without changing the bitstream.
|
|
|
|
Profiles are derived from a bandwidth figure, not chosen by eye:
|
|
|
|
```
|
|
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
|
|
```
|
|
|
|
> **On reading `docs/FINDINGS.md`:** it is append-only and several later sections
|
|
> overturn earlier ones. Superseded sections carry a blockquote at the top
|
|
> pointing to the correction — heed those, especially 18 (reversed by 21).
|
|
|
|
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored — supply your own.
|
|
|
|
**Not every large stream is game footage.** `00216` is the feature with a
|
|
burned-in commentary picture-in-picture and `00215` is the commentary itself —
|
|
the two largest files on the disc. The clean 9.4-minute animation is **`00223`**.
|
|
See FINDINGS 25.1 before running any size-ranked survey.
|