ROADMAP P5. The loader moved in session 21 and the frame clock in 22; the ring producer was the last policy living outside the machine. src/player/ring.i does `aligned` placement, the descriptor ring, a prefill, 51.2's slack rule and a seek, and the host keeps only the transport. It needed a container change. `aligned` asks whether the next record fits before the end of the ring -- a length asked BEFORE the record is fetched -- and every reader in this tree answered that by walking the frame stream, which is exactly what a player streaming off a disc cannot do. DLX4 carries nframes u16 record lengths in the scene header. Frame payloads are byte-identical to the DLX3 encode, so no fitted constant moves; the scene header goes 5,920 to 6,164 B. The producer reproduces the host's tiling exactly: 18 wraps, 14.7 KB mean hole, pixel-exact, a third independent implementation of the same policy. What it exposed is bigger than the item. A channel only moves bytes while it has a request and only the CPU can issue one, so the disc stands still between records by an amount the PLAYER sets, not the medium -- and no host-filled run could see it. At 488 KB/s in a 256 KB ring a one-deep request queue gives away 6.8% of the pipe and underruns 59 of 120 frames; two-deep gives away 3.4% and underruns none. The container's whole surplus over the wire is 8.7%, so the player's own loop was spending most of the slack a branch point saves up. Prefill is the weaker lever: six records of it still leaves 24 underruns. Three silent bugs are recorded in FINDINGS 55.7 -- all produced wrong pixels or a desync rather than a fault -- plus a rig one: MAME renders a screen line by line, so snapshotting the frame the decoder finished in captures a tear that reads exactly like a decoder bug. check.sh gains the machine-owned ring and a seek with the decode after it. decode.bin is unchanged at 1,296 B and a host-filled run executes none of the new code, so every FINDINGS 49/51 figure stands. ALL GREEN before and after. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
355 lines
20 KiB
Markdown
355 lines
20 KiB
Markdown
# Dragon's Lair: Sharp X68000 port
|
|
|
|
Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SCSI).
|
|
|
|
This is fundamentally a **video codec problem**, not a game-logic problem. The
|
|
game logic is a scene table with branching input windows; the difficulty is
|
|
pushing ~22 minutes of Don Bluth animation through a 10MHz 68000.
|
|
|
|
## What it looks like
|
|
|
|

|
|
|
|
Left, the Blu-ray frame cropped to 256x192. Right, the same frame **as the
|
|
emulated 68000 actually drew it**: 256 colours out of the X68000's 65536, one
|
|
16-colour-per-4x4-block codebook, decoded by `src/player/decode.s` from the
|
|
container. Not a re-render. These are the pixels MAME had on screen, pulled out
|
|
of its own snapshot, 2x nearest-neighbour, no filtering.
|
|
|
|
**The player, running.** 119 frames out of a **256 KB ring buffer on an emulated
|
|
stock 2 MB X68000**, paced to a 12 fps frame clock, streamed from a host file at
|
|
488 KB/s by `src/player/stream.s` with no Lua in the decode path. Source on the
|
|
left, the machine's screen on the right. (This recording was paced by the host;
|
|
the 68000 now keeps that clock itself, off the CRTC's V-DISP, and the same 120
|
|
frames decode pixel-exact under it — `src/player/clock.i`, FINDINGS 54.)
|
|
|
|
<video src="docs/img/player.webm" controls muted loop width="100%"></video>
|
|
|
|
[`docs/img/player.webm`](docs/img/player.webm) (119 frames, 12 fps, VP9)
|
|
|
|
116 of those 119 frames are **pixel-exact** against `tools/encoder/dlx.py`'s
|
|
reference reconstruction. The other three are **torn**: the top of the picture
|
|
is frame *n* and the bottom still holds frame *n-1*, because MAME captured the
|
|
screen while the block loop was partway down it. That is not a rig artefact.
|
|
`decode.s` writes straight to the displayed page, so a real player tears the
|
|
same way. `tools/media/make_readme_media.py` asserts the tear rather than
|
|
trimming it: every differing pixel has to come from the previous frame, or it
|
|
refuses to build.
|
|
|
|
**What the decoder is doing.** The same window with the block-mode map beside
|
|
it. **Black is SKIP** (costs nothing, draws nothing, the previous frame stands),
|
|
**blue is V1** (one codebook index for a whole 4x4 block), **amber is V4** (four
|
|
indices), **red is RAW** (sixteen bytes verbatim). The mode mix is what every
|
|
cost table in `docs/FINDINGS.md` is really about: V4 costs 1.5x V1, and the mode
|
|
decision is charged both bytes *and* cycles, which is why a byte-rich profile
|
|
buys its way out to RAW rather than V4.
|
|
|
|
<video src="docs/img/modes.webm" controls muted loop width="100%"></video>
|
|
|
|
[`docs/img/modes.webm`](docs/img/modes.webm) (the same 119 frames, with the mode map)
|
|
|
|
**Name the layer.** Everything above is **emulated**: MAME 0.277 `x68000`,
|
|
`-bios ipl10`, stock 10 MHz / 2 MB, cross-checked frame for frame on a second
|
|
CPU core (px68k's C68K). Nothing in this project has run on real hardware yet.
|
|
|
|
## Where it stands
|
|
|
|
**The binding resource is the 68000's local BUS, not its clock.** The decoder
|
|
occupies 86.7% of it once instruction prefetch is counted, and 52 of the 53
|
|
frames that miss the 12fps budget miss on the bus (FINDINGS 38). Read that
|
|
before optimising anything for cycles.
|
|
|
|
**The decoder works and is measured.** `decode.s` draws blocks and v7 literal
|
|
spans pixel-exact under both CPU cores, and costs inside the player what the
|
|
standalone blit benchmark said it would, to 0.2% (FINDINGS 41).
|
|
|
|
**The delivery path works too.** `stream.s` decodes the whole 120-frame window
|
|
out of a 256 KB ring on a stock 2 MB machine, final frame pixel-exact, with the
|
|
container in a host file rather than preloaded into RAM. The constraint is
|
|
**contiguity, not byte count**: the block loop reads with a monotonically
|
|
increasing `a0` and no bounds check, so the ring needs the whole next record
|
|
resident *and contiguous*, a condition no byte-counting buffer simulation can
|
|
see (FINDINGS 49).
|
|
|
|
**Seek slack is accumulated, not owned.** A ring's lookahead is built out of
|
|
`pipe - wire` and a seek spends all of it. At 488 KB/s a 256 KB ring needs 4.83
|
|
seconds of play to reach its 7-frame ceiling from empty, and 512 KB needs 8.42
|
|
seconds to reach 14, so a bigger ring raises the ceiling *and* lengthens the
|
|
climb. A branch point therefore asks "has there been enough play since the last
|
|
one", not "is the buffer big enough" (FINDINGS 51).
|
|
|
|
**There is no working delivery rate figure, deliberately.** `--bus`, `--kbps`
|
|
and `DLX_STREAM_KBPS` are required arguments with no defaults, so no table can
|
|
be scored against a rate its own output does not state. What replaces a constant
|
|
is a requirement: `tools/analysis/19_ring_stream.py` reports the **zero-prefill
|
|
pipe**, the rate a medium must clear for a container to need no prefill, which
|
|
is **513.2 KB/s** for the current candidate. That is a hardware acceptance test
|
|
to measure a BlueSCSI against (FINDINGS 50).
|
|
|
|
**The largest open number is W, the clocks stolen per delivered byte.** The
|
|
MB89352 is an 8-bit SPC, so the DMAC pays per byte rather than per word, which
|
|
is a 2x correction the project has already paid for once (FINDINGS 43). What W
|
|
costs is set by how the player programs the DMAC: 5 clocks a byte single
|
|
address with the bus held, 9 dual address held, 12 single address arbitrating
|
|
per byte, 16..19 dual address arbitrating per byte. The design's fate changes
|
|
completely across that ladder, and it is ours to choose.
|
|
|
|
**The one worked example on the machine is expensive.** The X68000 IPL ROM
|
|
programs all four HD63450 channels itself, and
|
|
`tools/analysis/21_iplrom_dmac.py` decodes that configuration out of the ROM
|
|
image and gates on the bytes still being there. Both the audio channel and the
|
|
on-board disk channel are dual address, 8-bit port, cycle steal *without* hold,
|
|
one external request per byte: **16..19 clocks a byte**, the top of the ladder.
|
|
For audio that is a settled figure and a small one, 1.25%..1.48% of a frame. For
|
|
the disk it is where nothing fits at any container size. The ROM drives SASI
|
|
rather than the MB89352, so it does not settle W, but a cheap configuration is
|
|
now the thing that has to be shown rather than assumed (FINDINGS 52).
|
|
|
|
**The player builds its own codebooks and palette now.** The two load-time
|
|
transforms — codebooks to word-per-pixel form, palette to `GGGGGRRRRRBBBBBI`
|
|
with the shared LSB picked per entry — ran host-side until session 21 and now
|
|
run on the 68000, out of the raw container header, byte-exact against the host
|
|
implementation on both CPU cores and with the palette read back out of the
|
|
hardware registers. A scene change costs **18.96 ms**, a third of one 12fps
|
|
frame slot. The finding underneath it is a cost nothing had counted: a scene
|
|
header is **5,920 bytes** that must arrive before frame 0, and in the currency
|
|
of seek slack those bytes lengthen the refill climb by 138 ms at 488 KB/s and by
|
|
**1.099 s at 451.4 KB/s**, because the surplus they are divided by goes to zero
|
|
(FINDINGS 53).
|
|
|
|
**The 68000 fills its own ring now, and the player's request loop costs more
|
|
than the medium does.** `src/player/ring.i` places records, prefills, keeps the
|
|
slack rule and seeks, out of a per-record index the container carries (DLX4).
|
|
The channel only moves bytes while it has a request and only the CPU can issue
|
|
one, so the disc **stands still between records** by an amount the player sets:
|
|
at 488 KB/s a one-deep request queue gives away **6.8% of the pipe and underruns
|
|
59 of 120 frames**, a two-deep one gives away 3.4% and underruns none — on a
|
|
container whose whole surplus over the wire is 8.7% (FINDINGS 55).
|
|
|
|
**Current encode:** 496.7 KB/s at 29.19 dB, 1 frame of 120 over the 12fps
|
|
budget, and that one is frame 0, the intra frame, late on purpose.
|
|
|
|
**Green-light check:** `./tools/bench/check.sh` (~4 min, needs the Blu-ray
|
|
mounted) re-runs both display regression tests, the rate-control drift gate, the
|
|
display-path coherency counterexample, a 120-frame 68000 decode on two CPU
|
|
cores, the ring and paced-ring passes, the DMAC configuration gate and the
|
|
load-time transforms on both cores, then prints `ALL GREEN`.
|
|
|
|
## Reproducing this
|
|
|
|
**No media ships in this repo and none of it is redistributable.** Bring your
|
|
own Dragon's Lair Blu-ray. Everything else needed to rebuild every number and
|
|
every picture above is either here or is packaged.
|
|
|
|
You need:
|
|
|
|
| | |
|
|
|---|---|
|
|
| the disc | loop-mounted read-only: `udisksctl loop-setup -r -f DRAGONS_LAIR.iso`. The tree was built against a decrypted UDF 2.x image. 7-Zip cannot read UDF 2.x, so use the loop mount |
|
|
| `python3` | plus **numpy** and **Pillow**, and nothing else. The k-means is hand-rolled rather than pulling in sklearn |
|
|
| `ffmpeg` / `ffprobe` | frame extraction, and the clips above |
|
|
| **MAME** | tested on 0.277, with the `x68000` ROM set. The rigs drive it headless via `-autoboot_script` |
|
|
| vasm (m68k, Motorola syntax) | **vendored**: `tools/vasm/vasmm68k_mot` is a Linux x86-64 binary, with the source tarball beside it to rebuild elsewhere |
|
|
|
|
Then:
|
|
|
|
```sh
|
|
export DLX_BDROM=/path/to/your/mounted/bluray # if not /media/$USER/BDROM
|
|
./tools/bench/check.sh # ~3 min, prints ALL GREEN
|
|
```
|
|
|
|
`DLX_BDROM` is honoured by every tool that reads the disc. Two stages are
|
|
optional and **skip rather than fail** when their input is absent, because both
|
|
live outside this repo:
|
|
|
|
- `PX68K=/path/to/px68k` for the second-CPU-core gate. This is the cheapest
|
|
strong test in the tree (seconds, no MAME, no ROMs) and it is what licenses
|
|
the bus and cycle figures.
|
|
- `IPLROM=/path/to/iplrom.dat` for the DMAC configuration gate. Defaults to
|
|
`~/mame/roms/iplrom.dat`.
|
|
|
|
To rebuild the stills and clips in `docs/img/` you also need a paced recording
|
|
run; see the header of `tools/media/make_readme_media.py`.
|
|
|
|
**Scene selection is a hard-coded stream number, not a search.** The gates use
|
|
streams `00020` and `00223` of the disc's 224 `.m2ts` files. A different
|
|
pressing may number them differently, and if so the green light will extract the
|
|
wrong footage rather than fail, so check that `tmp/fr_singe/` looks like the
|
|
Singe encounter before trusting any figure.
|
|
|
|
**Not every large stream is game footage.** `00216` is the feature with a
|
|
burned-in commentary picture-in-picture and `00215` is the commentary itself,
|
|
the two largest files on the disc. The clean 9.4-minute animation is **`00223`**
|
|
(FINDINGS 25.1).
|
|
|
|
## Encoder
|
|
|
|
```
|
|
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
|
|
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile scsi --preview p.png
|
|
```
|
|
|
|
The codec is a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4
|
|
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
|
|
rate-distortion. The RAW escape means `lam=0` is pixel-exact against the
|
|
palettised frame, so the quality knob spans lossless to heavily compressed
|
|
without changing the bitstream.
|
|
|
|
**Two byte budgets, not one.** `--kbps` is the quality rate point and
|
|
`--span-kbps` is the ceiling the span pass may draw on. They are different
|
|
things: the profile is chosen, the pipe is hardware, and bytes between them buy
|
|
a better picture if spent on `lam`, the 68000's deadline if spent on spans, and
|
|
nothing if left unspent. Spans run before `mu` because a span pays in bytes and
|
|
`mu` pays in picture (FINDINGS 41.2).
|
|
|
|
**Two ceilings, on two different axes.** The second is the 68000's decode
|
|
budget: `mu` is bisected per frame against 833,333 cycles so the frame also
|
|
*decodes* in time, which takes the worst sustained window from 37 frames over
|
|
budget to 1, for 0.62 dB at `scsi` (FINDINGS 31). It is on by default and
|
|
`--no-cpu-fit` turns it off. Unlike bytes, cycles have no bucket: there is no
|
|
double buffer to decode ahead into, so it is a hard per-frame ceiling.
|
|
|
|
**One profile, `scsi`, at 280 KB/s.** The 110 KB/s `sasi` profile was dropped on
|
|
capacity rather than bandwidth, since a SASI volume is limited to 40 MB and the
|
|
game's 22.8 minutes is 146 MiB even at that rate (FINDINGS 32). The rate point
|
|
may return under another name once the delivery medium is settled, because a 1x
|
|
CD-ROM sustains ~150 KB/s and CD-ROM is the only period medium with the
|
|
capacity.
|
|
|
|
The profile bitrate is a **ceiling**: `lam` is bisected per frame under a leaky
|
|
bucket, so the profile's `lam` is a quality floor rather than a setting
|
|
(`--fixed-lam` opts out). At `--spans all` none of that binds, though. A
|
|
32-frame bucket emits the same container byte for byte as an 8-frame one and
|
|
`lam` never leaves its floor on any frame of the reference window, because the
|
|
rate is set by the span pass and by `mu` (FINDINGS 44.3). Two known unit
|
|
inconsistencies on that side are implemented and default off because they
|
|
measure as a wash: `--joint-decide` prices a byte at `lam + mu*c` rather than
|
|
`lam`, and `--joint-bucket` stops the bucket lending clocks it cannot repay.
|
|
|
|
An encode is ~95% k-means. A 120-frame window is ~29 s, of which ~22 s is
|
|
training the two codebooks.
|
|
|
|
Profiles are derived from a bandwidth figure rather than chosen by eye:
|
|
|
|
```
|
|
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
|
|
```
|
|
|
|
## Documentation
|
|
|
|
- **`docs/STATUS.md`** is the current state, working setup, blockers and next
|
|
steps. **Start here.** It also lists what has been explicitly abandoned, so
|
|
old ideas do not get re-proposed.
|
|
- **`docs/ROADMAP.md`** is the remaining work to a completion target, and which
|
|
milestone that target is. Read it with STATUS rather than instead of it:
|
|
STATUS holds the measurements, ROADMAP holds the shape and goes stale first.
|
|
- **`docs/FINDINGS.md`** is measured hardware facts, content statistics, the
|
|
codec decision, and a section on measurement traps that produced three
|
|
separate false results. Read §4 before trusting any pipeline number. It is
|
|
append-only and later sections overturn earlier ones; superseded sections
|
|
carry a blockquote pointing at the correction.
|
|
- **`docs/BENCHMARK.md`** is how to measure the storage subsystem, and why a
|
|
bandwidth figure out of MAME would be meaningless.
|
|
- **`docs/HARDWARE.md`** is the X68000 GVRAM/CRTC reference.
|
|
|
|
## Layout
|
|
|
|
```
|
|
docs/ findings, status, roadmap, hardware reference
|
|
docs/img/ the stills and clips above, built from a real emulated run
|
|
tools/analysis/ measurement scripts, numbered in the order they were written.
|
|
Run from the repo root; they import from tools/encoder/.
|
|
01 and 02 are marked BROKEN deliberately and kept as
|
|
regression references.
|
|
10 is a COUNTEREXAMPLE and exits non-zero by design: it
|
|
demonstrates that the two-display-path plan corrupts 70 of 120
|
|
frames, which is why decode.s has one display path.
|
|
15 measures how much of the 68000's local bus the decoder
|
|
occupies and exits non-zero if its derived model stops
|
|
matching the harness's measurement.
|
|
16 is the DLX3 span container round-trip gate: it encodes,
|
|
writes the container, reads it back with the reference decoder
|
|
and fails if a pixel differs, or if it emitted too few spans to
|
|
have tested anything.
|
|
19 models the ring's ADDRESSES rather than its occupancy,
|
|
because each record must be contiguous and not merely resident,
|
|
and reports the zero-prefill pipe.
|
|
20 is an independent Python re-derivation of the seek-slack
|
|
model, sharing no code with the Lua producer it checks.
|
|
21 decodes the IPL ROM's HD63450 configuration and gates on the
|
|
bytes being where it says they are.
|
|
22 prices a scene change: header bytes, load-time clocks and
|
|
what both cost in accumulated seek slack, across explicit
|
|
rates. Its cycle counts are PARSED out of the rig's log, not
|
|
pasted in, so they cannot go stale silently.
|
|
buscost.py is the shared bus-cycle table. The per-block
|
|
constants live in tools/encoder/vq_hybrid.py and are imported,
|
|
never copied.
|
|
tools/bench/ MAME Lua injection harness and 68000 benchmark sources.
|
|
check.sh is the green light.
|
|
blit.s/blit.lua time the full-frame GVRAM blit on the 68000
|
|
itself. Not part of check.sh, because wall timings would make
|
|
the green light host-sensitive.
|
|
span.sh measures the literal-span mode the same way and
|
|
asserts that every one of its 36 timing configs drew a
|
|
pixel-exact frame, the count taken from generated metadata so
|
|
a new config cannot weaken the gate.
|
|
crtc_mode.lua is the single source of truth for CRTC R00-R08
|
|
and R20. Do not write CRTC values anywhere else.
|
|
prep_dlx.py/decode.lua/verify_decode.py load, time and verify
|
|
decode.s. prep_stream.py/stream.lua do the same for stream.s,
|
|
but lay the container out as a DISK in a host file and feed it
|
|
through a bounded ring at a modelled pipe rate, so the rig is
|
|
not bounded by the emulated machine's RAM and a stock 2 MB
|
|
machine runs the whole window. dlxload.py holds the
|
|
codebook/palette load-time maths both preps share -- and
|
|
the reference src/player/load.i is gated against.
|
|
prep_load.py/load.lua/verify_load.py/load_run.sh run those
|
|
transforms ON the 68000 and compare all 10,752 output bytes
|
|
with dlxload.py's, palette words read back out of the palette
|
|
registers rather than a RAM shadow.
|
|
tools/bench/c68k/ headless px68k C68K harness, a SECOND emulator for every
|
|
68000 cycle figure. Links only px68k's CPU core: no SDL, no
|
|
ROMs, no emulated machine. `make PX68K=~/src/px68k` then
|
|
run.sh; verify_c68k.py checks the decode is pixel-exact, which
|
|
is what licenses the cycle numbers. It also counts BUS cycles,
|
|
which MAME cannot report. The Makefile's -no-pie and the
|
|
harness's MAP_32BIT arena are load-bearing: C68K truncates
|
|
host pointers to 32 bits.
|
|
tools/media/ builds docs/img/ from a paced recording run
|
|
tools/vasm/ vasm m68k assembler, binary plus source tarball
|
|
tools/encoder/ hybrid VQ encoder and DLX3 container writer.
|
|
spans.py is the v7 span geometry, selection and serialiser,
|
|
and the single place the chain layout is stated on the encoder
|
|
side. It must match blit.s and decode.s: 11 coarse units of
|
|
24 px, 11 fine of 2.
|
|
DLX2 4-byte-aligns every frame record, because an odd move.l
|
|
is an ADDRESS ERROR on a 68000, not a slow read.
|
|
dlx.py is the reference DECODER, ground truth for the 68000.
|
|
24 models the ring with the 68000 owning it: the request
|
|
queue, the poll-only-when-not-decoding rule and 54.4's frame
|
|
cadence, and reports the pipe the player's own loop gives away.
|
|
src/player/ decode.s is the 68000 DLX3 decoder with a preloaded-stream
|
|
front-end. stream.s is the same decoder behind a bounded ring.
|
|
load.i is the LOAD-time half: codebook expansion and palette
|
|
packing, out of the raw container header, with loadgate.s as
|
|
its rig front-end. Its three scratch tables describe the
|
|
machine rather than the scene, so they are a separate entry
|
|
point a player calls once at boot.
|
|
ring.i is the RING PRODUCER: `aligned` placement, the
|
|
descriptor ring, the prefill policy, 51.2's slack rule as
|
|
arithmetic (ring_may_seek) and a seek. It reads the DLX4 record
|
|
index because a player cannot learn a record's length by
|
|
walking a stream it has not fetched.
|
|
Both include frame.i (the block loop and span chain) and
|
|
geom.i (the constants), so there is exactly ONE copy of the
|
|
bytes every cycle constant is fitted to. The span pass is
|
|
blit.s v7 verbatim, the same instruction sequence the
|
|
66.0/9.143/9.978 clock fit was measured on, so do not tidy it.
|
|
check.sh asserts decode.s still assembles to the same 1,296
|
|
bytes.
|
|
assets/ extracted frames and audio (gitignored)
|
|
```
|
|
|
|
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored. Supply your own.
|