diff --git a/README.md b/README.md index 8f14149..53b8b3d 100644 --- a/README.md +++ b/README.md @@ -1,143 +1,130 @@ -# Dragon's Lair — Sharp X68000 port +# Dragon's Lair: Sharp X68000 port Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SCSI). -This is fundamentally a **video codec problem**, not a game-logic problem: the +This is fundamentally a **video codec problem**, not a game-logic problem. The game logic is a scene table with branching input windows; the difficulty is pushing ~22 minutes of Don Bluth animation through a 10MHz 68000. ---- - ## What it looks like ![Blu-ray source next to the 68000's output](docs/img/source-vs-decoded.png) -Left, the Blu-ray frame cropped to 256x192. Right, **the same frame as the -emulated 68000 actually drew it** — 256 colours out of the X68000's 65536, one +Left, the Blu-ray frame cropped to 256x192. Right, the same frame **as the +emulated 68000 actually drew it**: 256 colours out of the X68000's 65536, one 16-colour-per-4x4-block codebook, decoded by `src/player/decode.s` from the -container. Not a re-render: these are the pixels MAME had on screen, extracted -from its own snapshot. 2x nearest-neighbour, no filtering. +container. Not a re-render. These are the pixels MAME had on screen, pulled out +of its own snapshot, 2x nearest-neighbour, no filtering. **The player, running.** 119 frames out of a **256 KB ring buffer on an emulated stock 2 MB X68000**, paced to a 12 fps frame clock, streamed from a host file at -488 KB/s — `src/player/stream.s`, no Lua in the decode path. Source on the left, -the machine's screen on the right. +488 KB/s by `src/player/stream.s` with no Lua in the decode path. Source on the +left, the machine's screen on the right. -[`docs/img/player.webm`](docs/img/player.webm) — 119 frames, 12 fps, VP9 +[`docs/img/player.webm`](docs/img/player.webm) (119 frames, 12 fps, VP9) 116 of those 119 frames are **pixel-exact** against `tools/encoder/dlx.py`'s -reference reconstruction. The other three are **torn** — the top of the picture +reference reconstruction. The other three are **torn**: the top of the picture is frame *n* and the bottom still holds frame *n-1*, because MAME captured the -screen while the block loop was partway down it. That is not a rig artefact: -`decode.s` writes straight to the displayed page (one display path, FINDINGS -28.1 — the dual-path plan is kept runnable as a counterexample precisely because -it corrupts frames), so a real player tears the same way. -`tools/media/make_readme_media.py` asserts the tear rather than trimming it — -every differing pixel has to come from the previous frame, or it refuses to -build. +screen while the block loop was partway down it. That is not a rig artefact. +`decode.s` writes straight to the displayed page, so a real player tears the +same way. `tools/media/make_readme_media.py` asserts the tear rather than +trimming it: every differing pixel has to come from the previous frame, or it +refuses to build. -**What the decoder is actually doing.** The same window with the block-mode map -beside it: **black = SKIP** (costs nothing, draws nothing — the previous frame -stands), **blue = V1** (one codebook index for a whole 4x4 block), **amber = V4** -(four indices), **red = RAW** (sixteen bytes verbatim). The mode mix is what -every cost table in `docs/FINDINGS.md` is really about — V4 costs 1.5x V1, and -since session 8 the mode decision is charged both bytes *and* cycles, which is -why a byte-rich profile buys its way out to RAW instead of V4. +**What the decoder is doing.** The same window with the block-mode map beside +it. **Black is SKIP** (costs nothing, draws nothing, the previous frame stands), +**blue is V1** (one codebook index for a whole 4x4 block), **amber is V4** (four +indices), **red is RAW** (sixteen bytes verbatim). The mode mix is what every +cost table in `docs/FINDINGS.md` is really about: V4 costs 1.5x V1, and the mode +decision is charged both bytes *and* cycles, which is why a byte-rich profile +buys its way out to RAW rather than V4. -[`docs/img/modes.webm`](docs/img/modes.webm) — the same 119 frames with the mode map +[`docs/img/modes.webm`](docs/img/modes.webm) (the same 119 frames, with the mode map) -**Name the layer:** everything above is **emulated** (MAME 0.277 `x68000`, -`-bios ipl10`, stock 10 MHz / 2 MB), cross-checked frame-for-frame on a second +**Name the layer.** Everything above is **emulated**: MAME 0.277 `x68000`, +`-bios ipl10`, stock 10 MHz / 2 MB, cross-checked frame for frame on a second CPU core (px68k's C68K). Nothing in this project has run on real hardware yet. ---- +## Where it stands -**And the binding resource is the 68000's local BUS, not its clock.** The -decoder occupies 86.7% of it once instruction prefetch is counted, and 52 of the -53 frames that miss the 12fps budget miss it on the bus, not the CPU -(FINDINGS 38). Read that before optimising anything for cycles. +**The binding resource is the 68000's local BUS, not its clock.** The decoder +occupies 86.7% of it once instruction prefetch is counted, and 52 of the 53 +frames that miss the 12fps budget miss on the bus (FINDINGS 38). Read that +before optimising anything for cycles. -The **literal span with a fine tail** (v7) is now IN the player: `decode.s` -paints it, pixel-exact under both CPU cores, and it costs inside the decoder -what `blit.s` said it would to 0.2% (FINDINGS 41). +**The decoder works and is measured.** `decode.s` draws blocks and v7 literal +spans pixel-exact under both CPU cores, and costs inside the player what the +standalone blit benchmark said it would, to 0.2% (FINDINGS 41). -**The delivery path is built and tested too** (FINDINGS 49). `src/player/stream.s` -decodes the whole 120-frame window **out of a 256 KB ring on a stock 2 MB -machine**, final frame pixel-exact, with the container in a host file rather than -preloaded into RAM. The constraint turned out to be **contiguity, not byte -count** — the block loop reads with a monotonically increasing `a0` and no bounds -check, so the ring needs the whole next record resident *and contiguous*, which -is a condition no byte-counting buffer simulation can see. +**The delivery path works too.** `stream.s` decodes the whole 120-frame window +out of a 256 KB ring on a stock 2 MB machine, final frame pixel-exact, with the +container in a host file rather than preloaded into RAM. The constraint is +**contiguity, not byte count**: the block loop reads with a monotonically +increasing `a0` and no bounds check, so the ring needs the whole next record +resident *and contiguous*, a condition no byte-counting buffer simulation can +see (FINDINGS 49). -**And building it caught a live defect**, then cost the project a constant. -The shipping candidate is 496.7 KB/s; the pipe figure the design had been -simulated against since session 2 was smaller, and nothing in the tree was -comparing the two — the rate controller binds on clocks and has no pipe term at -all, while the buffer sizing kept standing on a constant the design had stopped -enforcing. +**Seek slack is accumulated, not owned.** A ring's lookahead is built out of +`pipe - wire` and a seek spends all of it. At 488 KB/s a 256 KB ring needs 4.83 +seconds of play to reach its 7-frame ceiling from empty, and 512 KB needs 8.42 +seconds to reach 14, so a bigger ring raises the ceiling *and* lengthens the +climb. A branch point therefore asks "has there been enough play since the last +one", not "is the buffer big enough" (FINDINGS 51). -**So the pipe constant is retired (session 18, USER DECISION).** It was never a -bus measurement — user-supplied, no provenance, 10% of SCSI-1's asynchronous -rating (FINDINGS 42.1). It is gone as a default from every analysis tool and -from `stream.lua`; `--bus` / `--kbps` / `DLX_STREAM_KBPS` are now **required -arguments**, so no table can be scored against a rate its own output does not -state. **There is no working delivery figure, and that is the honest state.** +**There is no working delivery rate figure, deliberately.** `--bus`, `--kbps` +and `DLX_STREAM_KBPS` are required arguments with no defaults, so no table can +be scored against a rate its own output does not state. What replaces a constant +is a requirement: `tools/analysis/19_ring_stream.py` reports the **zero-prefill +pipe**, the rate a medium must clear for a container to need no prefill, which +is **513.2 KB/s** for the current candidate. That is a hardware acceptance test +to measure a BlueSCSI against (FINDINGS 50). -What replaces it is a requirement rather than a constant: -`tools/analysis/19_ring_stream.py` reports the **zero-prefill pipe**, the rate a -medium must clear for a container to need no prefill. For the candidate that is -**513.2 KB/s** — a hardware acceptance test to measure a BlueSCSI against. +**The largest open number is W, the clocks stolen per delivered byte.** The +MB89352 is an 8-bit SPC, so the DMAC pays per byte rather than per word, which +is a 2x correction the project has already paid for once (FINDINGS 43). What W +costs is set by how the player programs the DMAC: 5 clocks a byte single +address with the bus held, 9 dual address held, 12 single address arbitrating +per byte, 16..19 dual address arbitrating per byte. The design's fate changes +completely across that ladder, and it is ours to choose. -**Bytes are not free, and the number that said they were was in the wrong -unit.** Session 13 found the pipe figure the design was built against was never -a bus figure (SCSI-1 is 1.5 MB/s asynchronous) and concluded the span pass -saturates at ~837 KB/s, 0/120 frames over budget. Session 14 found the disk -debit behind that was charged **per word of stream to a byte-wide port** — the -MB89352 is an 8-bit SPC, so the DMAC pays per BYTE, and the debit is 2x every -table since FINDINGS 5. No 68000 bus cycle is shorter than four clocks, so the -old figure was below a physical floor. +**The one worked example on the machine is expensive.** The X68000 IPL ROM +programs all four HD63450 channels itself, and +`tools/analysis/21_iplrom_dmac.py` decodes that configuration out of the ROM +image and gates on the bytes still being there. Both the audio channel and the +on-board disk channel are dual address, 8-bit port, cycle steal *without* hold, +one external request per byte: **16..19 clocks a byte**, the top of the ladder. +For audio that is a settled figure and a small one, 1.25%..1.48% of a frame. For +the disk it is where nothing fits at any container size. The ROM drives SASI +rather than the MB89352, so it does not settle W, but a cheap configuration is +now the thing that has to be shown rather than assumed (FINDINGS 52). -**What survives, re-encoded honestly: 496.7 KB/s at 29.19 dB, 1 frame of 120 -over the 12fps budget** — and that one is frame 0, the intra frame, late on -purpose. The remaining lever is not ours: **whether the CZ-6BS1 wires the SPC's -DACK to the bus's `#EXACK`**, which decides 5 clocks/byte against 9, and with -it 242 KB/s and 0.69 dB. Read FINDINGS 43 before quoting any rate figure. - -**And session 20 read the answer the machine already had.** The X68000's IPL ROM -programs all four HD63450 channels itself, and MAME boots the rig with it, so -`tools/analysis/21_iplrom_dmac.py` decodes the configuration straight out of the -image and gates on the bytes still being there. The audio channel is -dual-address, 8-bit port, cycle steal **without hold**, one external request per -byte: **16..19 clocks a byte, not 5** — which prices the ADPCM stream at -1.25%..1.48% of a frame and closes ROADMAP's "do this first" item. The disk -channel is programmed **identically**. That is 16..19 clocks per delivered byte, -above the whole 5..12 bracket the project costs the transport in, and at that -price nothing fits. It is SASI and not the MB89352, so it does not settle the -question — but **a cheap configuration is now the thing that has to be shown, -not the thing assumed.** FINDINGS 52. +**Current encode:** 496.7 KB/s at 29.19 dB, 1 frame of 120 over the 12fps +budget, and that one is frame 0, the intra frame, late on purpose. **Green-light check:** `./tools/bench/check.sh` (~3 min, needs the Blu-ray -mounted) re-runs both display regression tests, the rate-control drift test, the -display-path coherency counterexample and a 120-frame 68000 decode, then prints -`ALL GREEN`. +mounted) re-runs both display regression tests, the rate-control drift gate, the +display-path coherency counterexample, a 120-frame 68000 decode on two CPU +cores, the ring and paced-ring passes and the DMAC configuration gate, then +prints `ALL GREEN`. ## Reproducing this -**No media ships in this repo, and none of it is redistributable.** Bring your +**No media ships in this repo and none of it is redistributable.** Bring your own Dragon's Lair Blu-ray. Everything else needed to rebuild every number and -every picture above is here or is packaged. +every picture above is either here or is packaged. You need: | | | |---|---| -| the disc | loop-mounted read-only. `udisksctl loop-setup -r -f DRAGONS_LAIR.iso` — the tree was built against a decrypted UDF 2.x image (7-Zip cannot read UDF 2.x; use the loop mount) | -| `python3` | plus **numpy** and **Pillow**. Nothing else — the k-means is hand-rolled rather than pulling in sklearn | -| `ffmpeg` / `ffprobe` | frame extraction, and the README clips | +| the disc | loop-mounted read-only: `udisksctl loop-setup -r -f DRAGONS_LAIR.iso`. The tree was built against a decrypted UDF 2.x image. 7-Zip cannot read UDF 2.x, so use the loop mount | +| `python3` | plus **numpy** and **Pillow**, and nothing else. The k-means is hand-rolled rather than pulling in sklearn | +| `ffmpeg` / `ffprobe` | frame extraction, and the clips above | | **MAME** | tested on 0.277, with the `x68000` ROM set. The rigs drive it headless via `-autoboot_script` | | vasm (m68k, Motorola syntax) | **vendored**: `tools/vasm/vasmm68k_mot` is a Linux x86-64 binary, with the source tarball beside it to rebuild elsewhere | @@ -152,136 +139,25 @@ export DLX_BDROM=/path/to/your/mounted/bluray # if not /media/$USER/BDROM optional and **skip rather than fail** when their input is absent, because both live outside this repo: -- `PX68K=/path/to/px68k` — a px68k checkout, for the second-CPU-core gate. This - is the cheapest strong test in the tree (seconds, no MAME, no ROMs), and it is - what licenses the bus and cycle figures. -- `IPLROM=/path/to/iplrom.dat` — the X68000 IPL ROM, for the DMAC-configuration - gate (FINDINGS 52). Defaults to `~/mame/roms/iplrom.dat`. +- `PX68K=/path/to/px68k` for the second-CPU-core gate. This is the cheapest + strong test in the tree (seconds, no MAME, no ROMs) and it is what licenses + the bus and cycle figures. +- `IPLROM=/path/to/iplrom.dat` for the DMAC configuration gate. Defaults to + `~/mame/roms/iplrom.dat`. To rebuild the stills and clips in `docs/img/` you also need a paced recording -run — see the header of `tools/media/make_readme_media.py`. +run; see the header of `tools/media/make_readme_media.py`. -**Scene selection is a hard-coded stream number**, not a search: the gates use -stream `00020` and `00223` of the disc's 224 `.m2ts` files, which are the ones -FINDINGS §1 and §25 characterise. A different pressing may number them -differently, and if so the green light will extract the wrong footage rather -than fail — check that `tmp/fr_singe/` looks like the Singe encounter (which is -what the directory is named for) before trusting any figure. +**Scene selection is a hard-coded stream number, not a search.** The gates use +streams `00020` and `00223` of the disc's 224 `.m2ts` files. A different +pressing may number them differently, and if so the green light will extract the +wrong footage rather than fail, so check that `tmp/fr_singe/` looks like the +Singe encounter before trusting any figure. -## Read first -- **`docs/FINDINGS.md`** — measured hardware facts, content statistics, codec - decision, and a section on measurement traps that produced three separate - false results. Read §4 before trusting any pipeline number. -- **`docs/STATUS.md`** — current state, working setup, blockers, next steps. - **Start here.** It also lists what has been explicitly abandoned, so old ideas - do not get re-proposed. -- **`docs/ROADMAP.md`** — the remaining work to a completion target, and which - milestone that target is. Read it with STATUS, not instead of it: STATUS holds - the measurements, ROADMAP holds the shape and goes stale first. -- **`docs/BENCHMARK.md`** — how to measure the storage subsystem, and why a - bandwidth figure out of MAME would be meaningless. -- **`docs/HARDWARE.md`** — X68000 GVRAM/CRTC reference. - -## Layout -``` -docs/ findings, status, hardware reference -tools/analysis/ measurement scripts, numbered in the order they were written - (01/02 marked BROKEN deliberately, kept as regression refs). - Run from the repo root — they import from tools/encoder/. - 07 finds the hottest sustained window in a stream; 08 renders - source | decoded | block-mode map as .webm; 09 is the - rate-control drift gate (FINDINGS 26/27) and is part of - check.sh -- it exits non-zero if the encoder ever again - reports a reconstruction no decoder would produce. - 10 is a COUNTEREXAMPLE, and exits non-zero by design: it - demonstrates that the two-display-path plan of FINDINGS - 24.5/25.6 corrupts 70 of 120 frames (FINDINGS 28.1). - 11 scores a container against the MEASURED per-mode block - costs without needing MAME; 12 prices the literal-span mode of - FINDINGS 30 against those same mode maps, and prints whether a - scene cut still fits at 12fps; 13 measures what fitting the - CPU budget costs in dB (FINDINGS 31) and caches H.build so the - search loop is seconds, not minutes. - 14 prices the HD63450 array-chain against the v6 and v7 - spans (FINDINGS 39/40) and prints the sensitivity that decides - it -- v7 is measured, and takes 37 of the 43 frames the DMAC - would, so the DMAC stays dropped; - 15 measures how much of the 68000's LOCAL bus the decoder - occupies (FINDINGS 38) and exits non-zero if its derived - model stops matching the harness's measurement. - 16 is the DLX3 span container ROUND-TRIP gate (part of - check.sh): it encodes, writes the container, reads it back with - the reference decoder and fails if a pixel differs -- or if it - emitted too few spans to have tested anything. 17 prices the - spans the encoder ACTUALLY emitted, with no selection model, - which is what 12 and 14 could only simulate. - 18 measures what a 16-colour text-plane literal would cost in - dB, and closes that direction (FINDINGS 46.3). - 19 is the RING-BUFFER simulation, and it supersedes 09_buffer_ - sim.py's question rather than repeating it: it models the ring's - ADDRESSES, because src/player/ needs each record contiguous and - not merely resident. It reports the ZERO-PREFILL PIPE -- the - rate a medium must clear for a container to need no prefill -- - and warns explicitly when demand exceeds supply on the MEAN, - where a "required prefill" figure would flatter a sustained - overrun. Its wrap count and hole size match tools/bench/ - stream.lua's, measured on a real 68000, to the digit. - buscost.py is the shared bus-cycle table both import; the - per-BLOCK constants live in tools/encoder/vq_hybrid.py and are - imported, never copied (session 12 corrected one of them). -tools/bench/ MAME Lua injection harness + 68000 benchmark sources. - `check.sh` re-runs both display regression tests (~40 s). - `blit.s`/`blit.lua` time the full-frame GVRAM blit on the - 68000 itself (FINDINGS 24) — not part of check.sh, because - wall timings would make the green-light check host-sensitive. - `span.sh` (prep_spans.py + span.lua + blit.s v5/v6/v7) - measures the literal-span mode the same way (FINDINGS 30 and - 40, ~30 s); it also asserts that every one of its 36 timing - configs drew a pixel-exact frame, the count taken from the - generated metadata so a new config cannot weaken the gate. - v7 is v6 with a second, 2-pixel chain for the span tail: - 66.0 cycles/span + 9.143 per coarse pixel + 9.978 per fine - pixel, MEASURED, which is the win FINDINGS 39.4 predicted. - `crtc_mode.lua` is the single source of truth for CRTC R00-R08 - and R20 — do not write CRTC values anywhere else. - `prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and - verify `src/player/decode.s`; the verify pass is in check.sh. - `prep_stream.py`/`stream.lua` do the same for `stream.s`, but - lay the container out as a DISK in a host file and feed it - through a bounded ring at a modelled pipe rate -- so the rig is - no longer bounded by the emulated machine's RAM, and a stock - 2 MB machine runs the whole window. `dlxload.py` holds the - codebook/palette load-time maths both preps share. -tools/bench/c68k/ headless px68k C68K harness -- a SECOND emulator for every - 68000 cycle figure (FINDINGS 37). Links only px68k's CPU core: - no SDL, no ROMs, no emulated machine. `make PX68K=~/src/px68k` - then `run.sh`; `verify_c68k.py` checks the decode is - pixel-exact, which is what licenses the cycle numbers. It also - counts BUS cycles, which MAME cannot report. - The Makefile's -no-pie and the harness's MAP_32BIT arena are - load-bearing: C68K truncates host pointers to 32 bits. -tools/vasm/ vasm m68k assembler (built from source) -tools/encoder/ hybrid VQ encoder + DLX3 container writer (working). - spans.py is the v7 span geometry, selection and serialiser, and - the single place the chain layout is stated on the encoder side - -- it must match blit.s/decode.s (11 coarse units of 24 px, 11 - fine of 2). - DLX2 4-byte-aligns every frame record: an odd `move.l` is an - ADDRESS ERROR on a 68000, not a slow read (FINDINGS 28.3). - dlx.py is the reference DECODER -- ground truth for the 68000. -src/player/ decode.s: the 68000 DLX3 decoder, PRELOADED-stream front-end. - stream.s: the same decoder behind a bounded RING (FINDINGS 49). - Both include frame.i (the block loop and span chain) and geom.i - (the constants) so there is exactly ONE copy of the bytes every - cycle constant in FINDINGS 24/30/40/41 is fitted to. check.sh - asserts decode.s still assembles to the same 1,296 bytes. - decode.s: the 68000 DLX3 decoder. Pixel-exact under MAME and - px68k's C68K core, blocks and v7 literal spans both. The span - pass is blit.s v7 verbatim -- the same instruction sequence the - 66.0/9.143/9.978 fit was measured on, so do not tidy it. - See FINDINGS 28, 31, 40 and 41. -assets/ extracted frames/audio (gitignored) -``` +**Not every large stream is game footage.** `00216` is the feature with a +burned-in commentary picture-in-picture and `00215` is the commentary itself, +the two largest files on the disc. The clean 9.4-minute animation is **`00223`** +(FINDINGS 25.1). ## Encoder @@ -290,61 +166,144 @@ python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop python3 tools/encoder/encode.py /tmp/fr out.dlx --profile scsi --preview p.png ``` -**Two budgets, not one.** `--kbps` is the quality rate point and `--span-kbps` -is the ceiling the span pass may draw on. They are different things: the profile -is chosen, the pipe is hardware, and bytes between them buy a better picture if -spent on `lam`, the 68000's deadline if spent on spans, and nothing if left -unspent. Spans run before `mu` because a span pays in bytes and `mu` pays in -picture (FINDINGS 41.2). +The codec is a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 +codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by +rate-distortion. The RAW escape means `lam=0` is pixel-exact against the +palettised frame, so the quality knob spans lossless to heavily compressed +without changing the bitstream. -**One profile: `scsi`, 280 KB/s.** The 110 KB/s `sasi` profile was dropped in -session 9 on capacity, not bandwidth — a SASI volume is limited to 40 MB, and -the game's 22.8 minutes of footage is 146 MiB even at that rate (FINDINGS 32). -The rate point may return under another name once the delivery medium is -settled, because a 1x CD-ROM sustains ~150 KB/s and CD-ROM is the only period -medium with the capacity. +**Two byte budgets, not one.** `--kbps` is the quality rate point and +`--span-kbps` is the ceiling the span pass may draw on. They are different +things: the profile is chosen, the pipe is hardware, and bytes between them buy +a better picture if spent on `lam`, the 68000's deadline if spent on spans, and +nothing if left unspent. Spans run before `mu` because a span pays in bytes and +`mu` pays in picture (FINDINGS 41.2). -The profile bitrate is a **ceiling**: lam is bisected per frame under a leaky +**Two ceilings, on two different axes.** The second is the 68000's decode +budget: `mu` is bisected per frame against 833,333 cycles so the frame also +*decodes* in time, which takes the worst sustained window from 37 frames over +budget to 1, for 0.62 dB at `scsi` (FINDINGS 31). It is on by default and +`--no-cpu-fit` turns it off. Unlike bytes, cycles have no bucket: there is no +double buffer to decode ahead into, so it is a hard per-frame ceiling. + +**One profile, `scsi`, at 280 KB/s.** The 110 KB/s `sasi` profile was dropped on +capacity rather than bandwidth, since a SASI volume is limited to 40 MB and the +game's 22.8 minutes is 146 MiB even at that rate (FINDINGS 32). The rate point +may return under another name once the delivery medium is settled, because a 1x +CD-ROM sustains ~150 KB/s and CD-ROM is the only period medium with the +capacity. + +The profile bitrate is a **ceiling**: `lam` is bisected per frame under a leaky bucket, so the profile's `lam` is a quality floor rather than a setting -(`--fixed-lam` opts out). +(`--fixed-lam` opts out). At `--spans all` none of that binds, though. A +32-frame bucket emits the same container byte for byte as an 8-frame one and +`lam` never leaves its floor on any frame of the reference window, because the +rate is set by the span pass and by `mu` (FINDINGS 44.3). Two known unit +inconsistencies on that side are implemented and default off because they +measure as a wash: `--joint-decide` prices a byte at `lam + mu*c` rather than +`lam`, and `--joint-bucket` stops the bucket lending clocks it cannot repay. -**But at `--spans all` none of that binds.** A 32-frame bucket emits the same -container byte for byte as an 8-frame one, and `lam` never leaves its floor on -any frame of the reference window: the rate is set by the span pass and by `mu`, -not by `--kbps` or the bucket (FINDINGS 44.3). Two known unit inconsistencies on -that side are implemented and default OFF because they measure as a wash -- -`--joint-decide` (the per-block lagrangian prices a byte at `lam + mu*c` rather -than `lam`) and `--joint-bucket` (the bucket may not lend clocks it cannot -repay). FINDINGS 44. +An encode is ~95% k-means. A 120-frame window is ~29 s, of which ~22 s is +training the two codebooks. -An encode is ~95% k-means; a 120-frame window is ~29 s, of which ~22 s is -training the two codebooks (FINDINGS 44.5). - -There are **two** ceilings, on two different axes. The second is the 68000's -decode budget: `mu` is bisected per frame against 833,333 cycles so the frame -also *decodes* in time, which takes the worst sustained window from 37 frames -over budget to 1 for 0.62 dB at `scsi` (FINDINGS 31). It is on by default; `--no-cpu-fit` -restores session 7 behaviour. Unlike bytes, cycles have no bucket — there is no -double buffer to decode ahead into, so it is a hard per-frame ceiling. The codec is -a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four -2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion. - -The RAW escape means `lam=0` is pixel-exact against the palettised frame, so the -quality knob spans lossless to heavily-compressed without changing the bitstream. - -Profiles are derived from a bandwidth figure, not chosen by eye: +Profiles are derived from a bandwidth figure rather than chosen by eye: ``` python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi ``` -> **On reading `docs/FINDINGS.md`:** it is append-only and several later sections -> overturn earlier ones. Superseded sections carry a blockquote at the top -> pointing to the correction — heed those, especially 18 (reversed by 21). +## Documentation -Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored — supply your own. +- **`docs/STATUS.md`** is the current state, working setup, blockers and next + steps. **Start here.** It also lists what has been explicitly abandoned, so + old ideas do not get re-proposed. +- **`docs/ROADMAP.md`** is the remaining work to a completion target, and which + milestone that target is. Read it with STATUS rather than instead of it: + STATUS holds the measurements, ROADMAP holds the shape and goes stale first. +- **`docs/FINDINGS.md`** is measured hardware facts, content statistics, the + codec decision, and a section on measurement traps that produced three + separate false results. Read §4 before trusting any pipeline number. It is + append-only and later sections overturn earlier ones; superseded sections + carry a blockquote pointing at the correction. +- **`docs/BENCHMARK.md`** is how to measure the storage subsystem, and why a + bandwidth figure out of MAME would be meaningless. +- **`docs/HARDWARE.md`** is the X68000 GVRAM/CRTC reference. -**Not every large stream is game footage.** `00216` is the feature with a -burned-in commentary picture-in-picture and `00215` is the commentary itself — -the two largest files on the disc. The clean 9.4-minute animation is **`00223`**. -See FINDINGS 25.1 before running any size-ranked survey. +## Layout + +``` +docs/ findings, status, roadmap, hardware reference +docs/img/ the stills and clips above, built from a real emulated run +tools/analysis/ measurement scripts, numbered in the order they were written. + Run from the repo root; they import from tools/encoder/. + 01 and 02 are marked BROKEN deliberately and kept as + regression references. + 10 is a COUNTEREXAMPLE and exits non-zero by design: it + demonstrates that the two-display-path plan corrupts 70 of 120 + frames, which is why decode.s has one display path. + 15 measures how much of the 68000's local bus the decoder + occupies and exits non-zero if its derived model stops + matching the harness's measurement. + 16 is the DLX3 span container round-trip gate: it encodes, + writes the container, reads it back with the reference decoder + and fails if a pixel differs, or if it emitted too few spans to + have tested anything. + 19 models the ring's ADDRESSES rather than its occupancy, + because each record must be contiguous and not merely resident, + and reports the zero-prefill pipe. + 20 is an independent Python re-derivation of the seek-slack + model, sharing no code with the Lua producer it checks. + 21 decodes the IPL ROM's HD63450 configuration and gates on the + bytes being where it says they are. + buscost.py is the shared bus-cycle table. The per-block + constants live in tools/encoder/vq_hybrid.py and are imported, + never copied. +tools/bench/ MAME Lua injection harness and 68000 benchmark sources. + check.sh is the green light. + blit.s/blit.lua time the full-frame GVRAM blit on the 68000 + itself. Not part of check.sh, because wall timings would make + the green light host-sensitive. + span.sh measures the literal-span mode the same way and + asserts that every one of its 36 timing configs drew a + pixel-exact frame, the count taken from generated metadata so + a new config cannot weaken the gate. + crtc_mode.lua is the single source of truth for CRTC R00-R08 + and R20. Do not write CRTC values anywhere else. + prep_dlx.py/decode.lua/verify_decode.py load, time and verify + decode.s. prep_stream.py/stream.lua do the same for stream.s, + but lay the container out as a DISK in a host file and feed it + through a bounded ring at a modelled pipe rate, so the rig is + not bounded by the emulated machine's RAM and a stock 2 MB + machine runs the whole window. dlxload.py holds the + codebook/palette load-time maths both preps share. +tools/bench/c68k/ headless px68k C68K harness, a SECOND emulator for every + 68000 cycle figure. Links only px68k's CPU core: no SDL, no + ROMs, no emulated machine. `make PX68K=~/src/px68k` then + run.sh; verify_c68k.py checks the decode is pixel-exact, which + is what licenses the cycle numbers. It also counts BUS cycles, + which MAME cannot report. The Makefile's -no-pie and the + harness's MAP_32BIT arena are load-bearing: C68K truncates + host pointers to 32 bits. +tools/media/ builds docs/img/ from a paced recording run +tools/vasm/ vasm m68k assembler, binary plus source tarball +tools/encoder/ hybrid VQ encoder and DLX3 container writer. + spans.py is the v7 span geometry, selection and serialiser, + and the single place the chain layout is stated on the encoder + side. It must match blit.s and decode.s: 11 coarse units of + 24 px, 11 fine of 2. + DLX2 4-byte-aligns every frame record, because an odd move.l + is an ADDRESS ERROR on a 68000, not a slow read. + dlx.py is the reference DECODER, ground truth for the 68000. +src/player/ decode.s is the 68000 DLX3 decoder with a preloaded-stream + front-end. stream.s is the same decoder behind a bounded ring. + Both include frame.i (the block loop and span chain) and + geom.i (the constants), so there is exactly ONE copy of the + bytes every cycle constant is fitted to. The span pass is + blit.s v7 verbatim, the same instruction sequence the + 66.0/9.143/9.978 clock fit was measured on, so do not tidy it. + check.sh asserts decode.s still assembles to the same 1,296 + bytes. +assets/ extracted frames and audio (gitignored) +``` + +Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored. Supply your own. diff --git a/docs/FINDINGS.md b/docs/FINDINGS.md index 851abb9..cb97636 100644 --- a/docs/FINDINGS.md +++ b/docs/FINDINGS.md @@ -4407,11 +4407,27 @@ without hold, external request.** Byte by byte, full arbitration each time. ch0 identically, and by 52.2's arithmetic that is **16..19 clocks per delivered byte**. -FINDINGS 42.4–42.6 brackets W, the clocks stolen per delivered byte, at **5..12**, -and reports that **W ≤ 6 fits 0/120 frames while W = 8 misses 47/120**. The only -worked example of a disk DMA configuration on this machine sits **above the -entire bracket.** `15_bus_occupancy.py` now sweeps it on the gate container -(37,403 B/frame, measured mean decode 570,958 clocks): +**A CORRECTION TO THIS SECTION AS FIRST WRITTEN, made in the same session.** +It cited 42.4's sensitivity table — `W <= 6` fits 0/120 frames, `W = 8` misses +47/120 — as though those figures were in clocks per BYTE. **They are per WORD, +and FINDINGS 43 voided them**: 43 is the section that caught `W` being charged +per word to a byte-wide port, and it says in terms that 42.3's 0/120 was never +physically reachable. Quoting them here would have re-imported the exact 2x unit +error 43 exists to have corrected, one section after using the same trap as a +warning. They are struck, and nothing below depends on them. + +In the corrected unit the ladder is a per-byte cost of the DMAC's own +configuration, and it is the ladder `buscost.py` already carries: + +| configuration | clk/byte | +|---|---:| +| single address, bus held | 5 | +| dual address, bus held | 9 | +| single address, arbitrated per byte | 12 | +| **dual address, arbitrated per byte — what the ROM programs** | **16..19** | + +`15_bus_occupancy.py` sweeps it on the gate container (37,403 B/frame, measured +mean decode 570,958 clocks): | W clk/B | video clk/frame | % of frame | CPU + audio + video | |---:|---:|---:|---:| @@ -4421,18 +4437,21 @@ entire bracket.** `15_bus_occupancy.py` now sweeps it on the gate container | **16** | 598,455 | 71.8% | **141.6%** | | **19** | 710,666 | 85.3% | **155.0%** | -The W = 8 row agreeing with 42.5's "misses 47/120" is a cross-check, not a new -result — two models of the same machine, one built from mode histograms and one -from bus clocks, landing in the same place. +**This table is the statement, and it is not corroborated by 42.4.** The +resemblance between the `W = 8` row here and 42.4's 47/120 is a coincidence of +two different units on two different containers at two different rates, and +calling it a cross-check — as this section did when first written — was +manufacturing agreement out of a unit error. **This does not close B3.** `scsiexrom.bin` drives an MB89352, not the SASI port, and a different ROM may configure it differently. What changed is the prior and the framing: 42.6 says the handshake "is ours to choose, not to receive", and that is still true — but **nothing in this tree has shown a cheaper configuration is reachable for an explicitly-addressed 8-bit port, and -the vendor's own answer is the expensive one.** W ≤ 12 is a *requirement on the -player's DMAC programming*, not a range the hardware hands us. It is now the -largest open number in the project, ahead of the rate. +the vendor's own answer is the expensive one.** Holding the bus, which is what +separates 9 from 16..19, is a *requirement on the player's DMAC programming* +rather than a range the hardware hands us. It is now the largest open number in +the project, ahead of the rate. ### 52.6 Audio outranks the disk at the arbiter CPR: FDC 0, **ADPCM 1**, SASI 2, `_DMAMOVE` 3 — lower is higher priority. With diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 16c40a9..f809b39 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -27,7 +27,7 @@ these units: | **68000 clocks** | measured, and the rate controller binds on them. | | **Delivery rate** | **no working figure, deliberately** (FINDINGS 50, USER DECISION). Every tool REQUIRES an explicit rate. | | **Seek time** | **no figure at all, and never had one.** 51.3/51.4 made it matter. | -| **W, clocks stolen per delivered byte** | bracketed 5..12 (42.4); the IPL ROM's own disk channel is **16..19** (52.5). **The largest open number in the project.** | +| **W, clocks stolen per delivered byte** | 5 single-address held, 9 dual held, 12 single arbitrated; the IPL ROM's own disk channel is **16..19** (52.5). **The largest open number in the project.** | --- @@ -122,12 +122,18 @@ receive** (FINDINGS 42.4-42.6). B3 informs it. **Session 20 promoted this to the project's biggest open number.** FINDINGS 52.5 found the only worked example of a disk DMA configuration on this machine — the -IPL ROM's own — sitting at 16..19 clk/B, outside the bracket entirely, where the -whole design fails at any container size (`15_bus_occupancy.py` sweeps it). -`W <= 12` is now a **requirement on the player's DMAC programming**, not a range -the hardware hands us, and demonstrating a configuration that meets it is P4's +IPL ROM's own — sitting at **16..19 clk/B**, where the whole design fails at any +container size (`15_bus_occupancy.py` sweeps it). The per-byte ladder is 5 clk/B +single-address with the bus held, 9 dual-address held, 12 single-address +arbitrated, 16..19 dual-address arbitrated. **Getting the DMAC to hold the bus +is the difference between 9 and 19**, it is a property of how the player +programs the channel, and demonstrating a configuration that does it is P4's first job rather than its last. +**Do not quote 42.4's `W <= 6` / `W = 8` sensitivity table for this.** It is in +clocks per WORD and FINDINGS 43 voided it; 52.5 cited it in byte units when +first written and strikes it. + **P5. Seek and branch.** Per-record index (the `aligned` producer needs one anyway, 49.3), prefill policy, and the accumulated-slack rule from 51.3 made explicit in the player rather than implied by the rig.