ROADMAP K1, the packed player's one open structural item. A frame is a picture
AND a palette, and no run in this tree had pointed a DMA channel at the palette
registers. dmagate.s runs 7-9, gated by dma_run.sh and check.sh:
7. 512 B off the disc into $E82000, bus held -- byte-exact in 256 register
words, read back OUT OF the registers by the 68000;
8. the SAME transfer aimed at RAM -- byte-exact at $2C000, and 256 of 256
palette words still read the poison the CPU wrote, which is what attributes
run 7 to the channel's MAR rather than to the readback path;
9. ONE array-chained start across two kinds of destination -- the palette and
six picture rows at the 1,024 B line stride, 2,048 B byte-exact.
So a packed frame is one channel start: a 193-entry array, palette first, CPU
halted from the first byte to the last. The array is scene-constant, because
the packed layout spends both 256-colour pages and there is no page to flip.
What is left on the CPU per frame in the video path is the channel start and the
READ(10) -- no per-frame PAINT, which is not the same claim as no per-frame CPU.
The destination is POISONED first (62.1). Runs 4-6 wrote into RAM that was zero
and GVRAM that was stale against a record that is mostly pad; "it matches the
disc" was weaker than it read as. The host counts whether the poison actually
discriminates instead of assuming it: 511 of 512, and the gate refuses under 500.
And it opened a hardware item (62.4, ROADMAP B4). MAME maps the palette to
palette_device over memory_array, whose write16 is a plain COMBINE_DATA -- RAM
that honours mem_mask, with no handler that could refuse a byte write. Unlike
GVRAM's 256-colour arm there is nothing here to be wrong about, so the run
bounds the model and not the board. What a real palette register does with a
byte write is unmeasured. A negative costs 0.28% of a frame and nothing else.
29_packed_player.py now also prints the two rows with the per-frame palette
charged -- 55.7% of a frame on the chain, 582 KB/s -- alongside the picture-only
figures the codec comparison is quoted against.
check.sh ALL GREEN before (tmp/check_s30_start.log) and after
(tmp/check_s30_end.log).
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
522 lines
32 KiB
Markdown
522 lines
32 KiB
Markdown
# Dragon's Lair: Sharp X68000 port
|
||
|
||
Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SCSI).
|
||
|
||
This is fundamentally a **video codec problem**, not a game-logic problem. The
|
||
game logic is a scene table with branching input windows; the difficulty is
|
||
pushing ~22 minutes of Don Bluth animation through a 10MHz 68000.
|
||
|
||
## What it looks like
|
||
|
||

|
||
|
||
Left, the Blu-ray frame cropped to 256x192. Right, the same frame **as the
|
||
emulated 68000 actually drew it**: 256 colours out of the X68000's 65536, one
|
||
16-colour-per-4x4-block codebook, decoded by `src/player/decode.s` from the
|
||
container. Not a re-render. These are the pixels MAME had on screen, pulled out
|
||
of its own snapshot, 2x nearest-neighbour, no filtering.
|
||
|
||
**The player, running.** 119 frames out of a **256 KB ring buffer on an emulated
|
||
stock 2 MB X68000**, paced to a 12 fps frame clock, streamed from a host file at
|
||
488 KB/s by `src/player/stream.s` with no Lua in the decode path. Source on the
|
||
left, the machine's screen on the right. (This recording was paced by the host;
|
||
the 68000 now keeps that clock itself, off the CRTC's V-DISP, and the same 120
|
||
frames decode pixel-exact under it — `src/player/clock.i`, FINDINGS 54.)
|
||
|
||
<video src="docs/img/player.webm" controls muted loop width="100%"></video>
|
||
|
||
[`docs/img/player.webm`](docs/img/player.webm) (119 frames, 12 fps, VP9)
|
||
|
||
116 of those 119 frames are **pixel-exact** against `tools/encoder/dlx.py`'s
|
||
reference reconstruction. The other three are **torn**: the top of the picture
|
||
is frame *n* and the bottom still holds frame *n-1*, because MAME captured the
|
||
screen while the block loop was partway down it. That is not a rig artefact.
|
||
`decode.s` writes straight to the displayed page, so a real player tears the
|
||
same way. `tools/media/make_readme_media.py` asserts the tear rather than
|
||
trimming it: every differing pixel has to come from the previous frame, or it
|
||
refuses to build.
|
||
|
||
**What the decoder is doing.** The same window with the block-mode map beside
|
||
it. **Black is SKIP** (costs nothing, draws nothing, the previous frame stands),
|
||
**blue is V1** (one codebook index for a whole 4x4 block), **amber is V4** (four
|
||
indices), **red is RAW** (sixteen bytes verbatim). The mode mix is what every
|
||
cost table in `docs/FINDINGS.md` is really about: V4 costs 1.5x V1, and the mode
|
||
decision is charged both bytes *and* cycles, which is why a byte-rich profile
|
||
buys its way out to RAW rather than V4.
|
||
|
||
<video src="docs/img/modes.webm" controls muted loop width="100%"></video>
|
||
|
||
[`docs/img/modes.webm`](docs/img/modes.webm) (the same 119 frames, with the mode map)
|
||
|
||
**Name the layer.** Everything above is **emulated**: MAME 0.277 `x68000`,
|
||
`-bios ipl10`, stock 10 MHz / 2 MB, cross-checked frame for frame on a second
|
||
CPU core (px68k's C68K). Nothing in this project has run on real hardware yet.
|
||
|
||
## Where it stands
|
||
|
||
**The binding resource is the 68000's local BUS, not its clock.** The decoder
|
||
occupies 86.7% of it once instruction prefetch is counted, and 52 of the 53
|
||
frames that miss the 12fps budget miss on the bus (FINDINGS 38). Read that
|
||
before optimising anything for cycles.
|
||
|
||
**The decoder works and is measured.** `decode.s` draws blocks and v7 literal
|
||
spans pixel-exact under both CPU cores, and costs inside the player what the
|
||
standalone blit benchmark said it would, to 0.2% (FINDINGS 41).
|
||
|
||
**The delivery path works too.** `stream.s` decodes the whole 120-frame window
|
||
out of a 256 KB ring on a stock 2 MB machine, final frame pixel-exact, with the
|
||
container in a host file rather than preloaded into RAM. The constraint is
|
||
**contiguity, not byte count**: the block loop reads with a monotonically
|
||
increasing `a0` and no bounds check, so the ring needs the whole next record
|
||
resident *and contiguous*, a condition no byte-counting buffer simulation can
|
||
see (FINDINGS 49).
|
||
|
||
**Seek slack is accumulated, not owned.** A ring's lookahead is built out of
|
||
`pipe - wire` and a seek spends all of it. At 488 KB/s a 256 KB ring needs 4.83
|
||
seconds of play to reach its 7-frame ceiling from empty, and 512 KB needs 8.42
|
||
seconds to reach 14, so a bigger ring raises the ceiling *and* lengthens the
|
||
climb. A branch point therefore asks "has there been enough play since the last
|
||
one", not "is the buffer big enough" (FINDINGS 51).
|
||
|
||
**There is no working delivery rate figure, deliberately.** `--bus`, `--kbps`
|
||
and `DLX_STREAM_KBPS` are required arguments with no defaults, so no table can
|
||
be scored against a rate its own output does not state. What replaces a constant
|
||
is a requirement: `tools/analysis/19_ring_stream.py` reports the **zero-prefill
|
||
pipe**, the rate a medium must clear for a container to need no prefill, which
|
||
is **513.2 KB/s** for the current candidate. That is a hardware acceptance test
|
||
to measure a BlueSCSI against (FINDINGS 50).
|
||
|
||
**The largest open number is W, the clocks stolen per delivered byte.** The
|
||
MB89352 is an 8-bit SPC, so the DMAC pays per byte rather than per word, which
|
||
is a 2x correction the project has already paid for once (FINDINGS 43). What W
|
||
costs is set by how the player programs the DMAC: 5 clocks a byte single
|
||
address with the bus held, 9 dual address held, 12 single address arbitrating
|
||
per byte, 16..19 dual address arbitrating per byte. The design's fate changes
|
||
completely across that ladder, and it is ours to choose.
|
||
|
||
**The one worked example on the machine is expensive.** The X68000 IPL ROM
|
||
programs all four HD63450 channels itself, and
|
||
`tools/analysis/21_iplrom_dmac.py` decodes that configuration out of the ROM
|
||
image and gates on the bytes still being there. Both the audio channel and the
|
||
on-board disk channel are dual address, 8-bit port, cycle steal *without* hold,
|
||
one external request per byte: **16..19 clocks a byte**, the top of the ladder.
|
||
For audio that is a settled figure and a small one, 1.25%..1.48% of a frame. For
|
||
the disk it is where nothing fits at any container size. The ROM drives SASI
|
||
rather than the MB89352, so it does not settle W, but a cheap configuration is
|
||
now the thing that has to be shown rather than assumed (FINDINGS 52).
|
||
|
||
**The player builds its own codebooks and palette now.** The two load-time
|
||
transforms — codebooks to word-per-pixel form, palette to `GGGGGRRRRRBBBBBI`
|
||
with the shared LSB picked per entry — ran host-side until session 21 and now
|
||
run on the 68000, out of the raw container header, byte-exact against the host
|
||
implementation on both CPU cores and with the palette read back out of the
|
||
hardware registers. A scene change costs **18.96 ms**, a third of one 12fps
|
||
frame slot. The finding underneath it is a cost nothing had counted: a scene
|
||
header is **5,920 bytes** that must arrive before frame 0, and in the currency
|
||
of seek slack those bytes lengthen the refill climb by 138 ms at 488 KB/s and by
|
||
**1.099 s at 451.4 KB/s**, because the surplus they are divided by goes to zero
|
||
(FINDINGS 53).
|
||
|
||
**The 68000 fills its own ring now, and the player's request loop costs more
|
||
than the medium does.** `src/player/ring.i` places records, prefills, keeps the
|
||
slack rule and seeks, out of a per-record index the container carries (DLX4).
|
||
The channel only moves bytes while it has a request and only the CPU can issue
|
||
one, so the disc **stands still between records** by an amount the player sets:
|
||
at 488 KB/s a one-deep request queue gives away **6.8% of the pipe and underruns
|
||
59 of 120 frames**, a two-deep one gives away 3.4% and underruns none — on a
|
||
container whose whole surplus over the wire is 8.7% (FINDINGS 55).
|
||
|
||
**The player runs off a real disc now, and PIO costs 87 clocks a byte.**
|
||
`src/player/xfer.i` answers the ring's request mailbox with a real READ(10) to a
|
||
real MB89352 instead of a host moving bytes at a modelled rate: 120 records,
|
||
4,488,588 B, **pixel-exact out of a 256 KB ring**, with a real mid-stream seek in
|
||
a second pass, and the **same 18 wraps** three different transports have now
|
||
produced. What it costs is the finding. Subtracting the same 120 frames run
|
||
twice gives **87.28 clocks per delivered byte**, and the 68000's own cycle table
|
||
for that loop says **87.15** — 0.2% apart, so the cost is the instruction stream
|
||
rather than the emulator's device model, and it is the first number this rig has
|
||
produced that a real board would also pay. At this container's mean record that
|
||
is **391.8% of a 12 fps frame**; the machine's own V-DISP clock agrees from the
|
||
other end at **2.57 fps**. Against the W ladder — 22.4% of a frame at 5 clocks a
|
||
byte, 85.3% at 19 — **the CPU doing the work itself is 4.6x the worst DMA
|
||
configuration this project has found and 17.5x the best.** Getting the DMAC to
|
||
hold the bus is no longer worth 9 against 19; it is worth 87 against either, and
|
||
it is the only thing left before a player (FINDINGS 58).
|
||
|
||
**The DMAC drives the data phase now, and it holds the bus.** `src/player/dma.i`
|
||
programs an HD63450 channel and hands it the SCSI data phase: **the same 2,048
|
||
bytes come off the disc three ways — PIO, the channel with the bus held, the
|
||
channel stealing cycles — and all three are byte-exact.** The evidence that the
|
||
DMAC and not the CPU is driving it never looks at the data register, which
|
||
cannot answer the question: with the DMAC's OWN asserted, MAME cannot tell a
|
||
CPU-driven byte at `$EA0015` from a DMAC-driven one. What it looks at instead is
|
||
**the CPU's own progress**. MTC is sampled by the instruction *after* the one
|
||
that starts the channel; held, it reads **zero of 2,048** — the whole transfer
|
||
happened between two instructions, because the 68000 did not execute in between
|
||
— while the stealing configuration reads the full count and the CPU then goes
|
||
round its own loop 426 times. Put the stealing registers in the held slot and
|
||
the run still delivers every byte and the gate goes **red**, which is what says
|
||
the counter can come out different (FINDINGS 59.1).
|
||
|
||
**And auto-request is charged by time, not by byte.** The card as MAME models it
|
||
has **no request line to the DMAC at all** — its flow control is DTACK — so
|
||
every configuration that can be run against it is auto-request, and an
|
||
auto-requested channel does not know whether the device is ready: it spends its
|
||
share of the bus either way. Every `W` in this project is clocks per *delivered*
|
||
byte, which presumes the device asks; here the cost scales with **how long the
|
||
record takes to arrive**, so halving the delivery rate *doubles* the CPU cost of
|
||
the same record. Priced from the MC68450's own limited-rate constants against an
|
||
explicit 460 KB/s: max rate costs the whole **95.3% of a frame** the record takes
|
||
to land, and of the four bus shares the GCR can be programmed for — 50, 25,
|
||
12.5, 6.25% — **only 50% carries the rate**, at 10.61 clocks a byte and 47.6% of
|
||
a frame. The GCR is a design lever nothing in this tree had named (FINDINGS
|
||
59.3).
|
||
|
||
**And what it all costs: the frame affords 6.74 clocks a byte, and a
|
||
dual-address byte is 9.** Putting the transport on the channel cuts it from
|
||
**391.7% of a 12 fps frame to 40..95%** — four to ten times, the largest
|
||
movement in this project's cost model since the decoder was written — and **it
|
||
still does not fit.** After the measured decode (68.5%) and the audio DMA
|
||
(1.25%), 30.2% of the frame is left, which at this container's 37,403 B record
|
||
is 6.74 clocks a byte; a dual-address byte is a 4-clock read of the device plus
|
||
a 5-clock write to memory, so **9 is a floor no bus share and no delivery rate
|
||
goes under**. Single address is 5 and fits at 92.2% with room to spare — and it
|
||
needs the device to ACK the DMAC, which needs a request line MAME does not
|
||
connect and the slot pinout does have. So the project's live question is now a
|
||
fact about a board: **does a real CZ-6BS1 drive `#EXREQ`?** If it does, the
|
||
design fits. If it does not, the container has to come down from 438 KB/s of
|
||
payload to **328** — which is an encoder target, entirely inside this project,
|
||
and measured against the heaviest container the encoder emits rather than
|
||
against a shipping one (FINDINGS 59.7).
|
||
|
||
**A record was not a sector, and the fix was a re-encode — it is done.** 117 of
|
||
120 records used to start part way into a 512 B block, and reading whole blocks
|
||
into the ring corrupts the neighbouring records rather than merely wasting bytes
|
||
— the block loop reads with no bounds check. PIO absorbed this for free by
|
||
simply not storing the bytes outside the window, a property that disappears the
|
||
moment a DMA channel takes over. Priced three ways: windowed PIO is +1.34% on
|
||
the wire and cannot be done by a channel at all; a bounce buffer is +1.34% and
|
||
**+5 clocks on every delivered byte**, 22.4% of a frame; sector-aligning records
|
||
in the container is **+0.43% and zero clocks** (FINDINGS 58.3). Session 27 made
|
||
it a precondition rather than a preference — the transport *refuses* a windowed
|
||
read when the data phase is the channel's (59.4) — and **session 28 met it: the
|
||
container is DLX5, every record is padded to 512 B and the frame stream starts
|
||
on a sector boundary. 120 of 120 records are aligned, the realised wire cost is
|
||
+0.48%, and the disc now moves exactly the records** — the bytes off the disc
|
||
and the bytes into the ring are the same number, which is what check.sh gates on
|
||
(FINDINGS 60.1).
|
||
|
||
**And the player that has no decoder at all fits the budget the codec misses.**
|
||
256-colour GVRAM throws away the high byte of every word a CPU writes, so a
|
||
picture byte normally costs two disc bytes — but CRTC R20 bit 11 turns the
|
||
masking off, and with the two 256-colour pages scrolled apart one word carries
|
||
two pixels (FINDINGS 46/47). Session 29 measured what that is worth. The packed
|
||
full-frame blit is **227,553 clocks, 27.3% of a 12 fps frame** — 51% of the
|
||
unpacked one, and the *same* as the unpacked path's write-only floor, so packing
|
||
buys back the whole of the source read. A DMA channel fills GVRAM in buffer mode
|
||
straight off the disc with the CPU halted, and **walks the 1,024-byte line stride
|
||
itself** through array chaining, so a frame is one channel start and not 192. At
|
||
the 9 clk/B dual-address floor — the only configuration this machine can be shown
|
||
to run — **the shipping codec is 110.4% of a frame and a decoder-free packed
|
||
player is 55.2%.** Decoding 37,585 bytes costs more than not decoding 49,152.
|
||
|
||
What it costs is the wire: **576 KB/s, fixed, with no lever** — a codec's bitrate
|
||
is adjustable and a literal frame's is geometry — against 327 KB/s for the codec
|
||
at the same floor. So the two open hardware facts changed character: **whether
|
||
the medium sustains 576 KB/s, and whether buffer mode blanks the layer while it
|
||
is being written, now decide which player exists** rather than how much headroom
|
||
one has. The codec cannot take the packing either way: writing 4×4 blocks a byte
|
||
at a time is **28% dearer** than the shipping shape, and pairing the blocks 128
|
||
columns apart to get the burst back drops SKIP from 66.3% of blocks to 46.1% of
|
||
pairs — about **+60% on the bytes**, against a target that needs them 35% lower
|
||
(FINDINGS 61).
|
||
|
||
**So encoder work is PARKED (USER DECISION, session 29).** Not because the codec
|
||
is wrong, but because its remaining path is a conjunction and the packed one is
|
||
not. The codec that exists is 440 KB/s and 110.4% of a frame; reaching E7's
|
||
327 KB/s needs a 35% byte reduction after two of its three levers were measured
|
||
and found inert (60.4, 60.5), and the reward on success is a design at ~100% of
|
||
the frame. The packed player is at 55.2% today. **The codec is kept on disk and
|
||
not built on**, because B2 is unanswered and 48.1's prior leans against packing —
|
||
if buffer mode blanks, it is the only thing left (48.3).
|
||
|
||
**And a frame is now one channel start.** Session 30 asked the packed player's
|
||
one open structural question: a frame is a picture *and* a palette, and nothing
|
||
had ever pointed a DMA channel at the palette registers. It writes them —
|
||
512 B off the disc byte-exact into 256 registers at `$E82000`, read back out of
|
||
the registers by the 68000 — and **one array-chained start crosses from those
|
||
registers into GVRAM**, which is the shape of a whole frame: a palette entry and
|
||
192 row entries, walked by the channel with the CPU halted throughout. The array
|
||
is scene-constant, because the packed layout spends both 256-colour pages and
|
||
there is no page to flip. What is left on the CPU per frame in the video path is
|
||
the channel start and the disc read; there is no per-frame *paint*. **What it
|
||
does not settle is the board** — MAME models the palette as plain
|
||
`COMBINE_DATA` storage with no handler that could refuse a byte write, so the
|
||
run bounds the model and not the hardware, and "does a real palette register
|
||
take a byte write" joins the hardware list as B4. A negative answer costs 0.28%
|
||
of a frame and nothing else (FINDINGS 62).
|
||
|
||
**The scene graph is in, and the worst gap between two decision points is
|
||
zero.** `tools/import/scenegraph.py` imports the arcade scene graph — 40 scenes,
|
||
516 sequences, 906 input windows — and 5.4% of the game's 612 branch transitions
|
||
open an input window on the first frame of a clip the disc *seeked to*, so two
|
||
seeks can fall back to back with no play between them. A rule of the form "has
|
||
there been enough play since the last branch" can therefore be answered no by
|
||
the **content**, not by the buffer. It does not break the design: a branch on an
|
||
empty ring costs the 2-record prefill, **149.7 ms at 488 KB/s**, not the climb.
|
||
What it removes is margin — at that rate in a 256 KB ring, **76% of this game's
|
||
branch points arrive before the ring has refilled**, and a 512 KB ring makes it
|
||
90%, because doubling the ceiling does not touch `pipe - wire` (FINDINGS 56).
|
||
|
||
**Nothing outside-derived is committed here.** The scene graph is not
|
||
redistributable from this tree; it is regenerated from a reader's own clones
|
||
into gitignored `tmp/`, and `tools/import/scenegraph.py` is the single file in
|
||
the repo coupled to those projects — everything downstream reads `DLXSCENE1`,
|
||
this project's own schema, with the sources' attribution carried in it.
|
||
DirkSimple is zlib (Ryan C. Gordon); the SNES chapter set is MIT (Chad
|
||
Doebelin) and, by its own README, *derived* from DirkSimple rather than an
|
||
independent transcription, which struck a cross-check this project had planned
|
||
on for eight sessions.
|
||
|
||
**Current encode:** 496.7 KB/s at 29.19 dB, 1 frame of 120 over the 12fps
|
||
budget, and that one is frame 0, the intra frame, late on purpose.
|
||
|
||
**Green-light check:** `./tools/bench/check.sh` (~4 min, needs the Blu-ray
|
||
mounted) re-runs both display regression tests, the rate-control drift gate, the
|
||
display-path coherency counterexample, a 120-frame 68000 decode on two CPU
|
||
cores, the ring and paced-ring passes, the DMAC configuration gate and the
|
||
load-time transforms on both cores, then imports and gates the scene graph
|
||
when a DirkSimple checkout is present, then prints `ALL GREEN`.
|
||
|
||
## Reproducing this
|
||
|
||
**No media ships in this repo and none of it is redistributable.** Bring your
|
||
own Dragon's Lair Blu-ray. Everything else needed to rebuild every number and
|
||
every picture above is either here or is packaged.
|
||
|
||
You need:
|
||
|
||
| | |
|
||
|---|---|
|
||
| the disc | loop-mounted read-only: `udisksctl loop-setup -r -f DRAGONS_LAIR.iso`. The tree was built against a decrypted UDF 2.x image. 7-Zip cannot read UDF 2.x, so use the loop mount |
|
||
| `python3` | plus **numpy** and **Pillow**, and nothing else. The k-means is hand-rolled rather than pulling in sklearn |
|
||
| `ffmpeg` / `ffprobe` | frame extraction, and the clips above |
|
||
| **MAME** | tested on 0.277, with the `x68000` ROM set. The rigs drive it headless via `-autoboot_script` |
|
||
| vasm (m68k, Motorola syntax) | **vendored**: `tools/vasm/vasmm68k_mot` is a Linux x86-64 binary, with the source tarball beside it to rebuild elsewhere |
|
||
|
||
Then:
|
||
|
||
```sh
|
||
export DLX_BDROM=/path/to/your/mounted/bluray # if not /media/$USER/BDROM
|
||
./tools/bench/check.sh # ~3 min, prints ALL GREEN
|
||
```
|
||
|
||
`DLX_BDROM` is honoured by every tool that reads the disc. Two stages are
|
||
optional and **skip rather than fail** when their input is absent, because both
|
||
live outside this repo:
|
||
|
||
- `PX68K=/path/to/px68k` for the second-CPU-core gate. This is the cheapest
|
||
strong test in the tree (seconds, no MAME, no ROMs) and it is what licenses
|
||
the bus and cycle figures.
|
||
- `IPLROM=/path/to/iplrom.dat` for the DMAC configuration gate. Defaults to
|
||
`~/mame/roms/iplrom.dat`.
|
||
|
||
To rebuild the stills and clips in `docs/img/` you also need a paced recording
|
||
run; see the header of `tools/media/make_readme_media.py`.
|
||
|
||
**Scene selection is a hard-coded stream number, not a search.** The gates use
|
||
streams `00020` and `00223` of the disc's 224 `.m2ts` files. A different
|
||
pressing may number them differently, and if so the green light will extract the
|
||
wrong footage rather than fail, so check that `tmp/fr_singe/` looks like the
|
||
Singe encounter before trusting any figure.
|
||
|
||
**Not every large stream is game footage.** `00216` is the feature with a
|
||
burned-in commentary picture-in-picture and `00215` is the commentary itself,
|
||
the two largest files on the disc. The clean 9.4-minute animation is **`00223`**
|
||
(FINDINGS 25.1).
|
||
|
||
## Encoder
|
||
|
||
```
|
||
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
|
||
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile scsi --preview p.png
|
||
```
|
||
|
||
The codec is a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4
|
||
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
|
||
rate-distortion. The RAW escape means `lam=0` is pixel-exact against the
|
||
palettised frame, so the quality knob spans lossless to heavily compressed
|
||
without changing the bitstream.
|
||
|
||
**Two byte budgets, not one.** `--kbps` is the quality rate point and
|
||
`--span-kbps` is the ceiling the span pass may draw on. They are different
|
||
things: the profile is chosen, the pipe is hardware, and bytes between them buy
|
||
a better picture if spent on `lam`, the 68000's deadline if spent on spans, and
|
||
nothing if left unspent. Spans run before `mu` because a span pays in bytes and
|
||
`mu` pays in picture (FINDINGS 41.2).
|
||
|
||
**Two ceilings, on two different axes.** The second is the 68000's decode
|
||
budget: `mu` is bisected per frame against 833,333 cycles so the frame also
|
||
*decodes* in time, which takes the worst sustained window from 37 frames over
|
||
budget to 1, for 0.62 dB at `scsi` (FINDINGS 31). It is on by default and
|
||
`--no-cpu-fit` turns it off. Unlike bytes, cycles have no bucket: there is no
|
||
double buffer to decode ahead into, so it is a hard per-frame ceiling.
|
||
|
||
**One profile, `scsi`, at 280 KB/s.** The 110 KB/s `sasi` profile was dropped on
|
||
capacity rather than bandwidth, since a SASI volume is limited to 40 MB and the
|
||
game's 22.8 minutes is 146 MiB even at that rate (FINDINGS 32). The rate point
|
||
may return under another name once the delivery medium is settled, because a 1x
|
||
CD-ROM sustains ~150 KB/s and CD-ROM is the only period medium with the
|
||
capacity.
|
||
|
||
The profile bitrate is a **ceiling**: `lam` is bisected per frame under a leaky
|
||
bucket, so the profile's `lam` is a quality floor rather than a setting
|
||
(`--fixed-lam` opts out). At `--spans all` none of that binds, though. A
|
||
32-frame bucket emits the same container byte for byte as an 8-frame one and
|
||
`lam` never leaves its floor on any frame of the reference window, because the
|
||
rate is set by the span pass and by `mu` (FINDINGS 44.3). Two known unit
|
||
inconsistencies on that side are implemented and default off because they
|
||
measure as a wash: `--joint-decide` prices a byte at `lam + mu*c` rather than
|
||
`lam`, and `--joint-bucket` stops the bucket lending clocks it cannot repay.
|
||
|
||
An encode is ~95% k-means. A 120-frame window is ~29 s, of which ~22 s is
|
||
training the two codebooks.
|
||
|
||
Profiles are derived from a bandwidth figure rather than chosen by eye:
|
||
|
||
```
|
||
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
|
||
```
|
||
|
||
## Documentation
|
||
|
||
- **`docs/STATUS.md`** is the current state, working setup, blockers and next
|
||
steps. **Start here.** It also lists what has been explicitly abandoned, so
|
||
old ideas do not get re-proposed.
|
||
- **`docs/ROADMAP.md`** is the remaining work to a completion target, and which
|
||
milestone that target is. Read it with STATUS rather than instead of it:
|
||
STATUS holds the measurements, ROADMAP holds the shape and goes stale first.
|
||
- **`docs/FINDINGS.md`** is measured hardware facts, content statistics, the
|
||
codec decision, and a section on measurement traps that produced three
|
||
separate false results. Read §4 before trusting any pipeline number. It is
|
||
append-only and later sections overturn earlier ones; superseded sections
|
||
carry a blockquote pointing at the correction.
|
||
- **`docs/BENCHMARK.md`** is how to measure the storage subsystem, and why a
|
||
bandwidth figure out of MAME would be meaningless.
|
||
- **`docs/HARDWARE.md`** is the X68000 GVRAM/CRTC reference.
|
||
|
||
## Layout
|
||
|
||
```
|
||
docs/ findings, status, roadmap, hardware reference
|
||
docs/img/ the stills and clips above, built from a real emulated run
|
||
tools/analysis/ measurement scripts, numbered in the order they were written.
|
||
Run from the repo root; they import from tools/encoder/.
|
||
01 and 02 are marked BROKEN deliberately and kept as
|
||
regression references.
|
||
10 is a COUNTEREXAMPLE and exits non-zero by design: it
|
||
demonstrates that the two-display-path plan corrupts 70 of 120
|
||
frames, which is why decode.s has one display path.
|
||
15 measures how much of the 68000's local bus the decoder
|
||
occupies and exits non-zero if its derived model stops
|
||
matching the harness's measurement.
|
||
16 is the DLX3 span container round-trip gate: it encodes,
|
||
writes the container, reads it back with the reference decoder
|
||
and fails if a pixel differs, or if it emitted too few spans to
|
||
have tested anything.
|
||
19 models the ring's ADDRESSES rather than its occupancy,
|
||
because each record must be contiguous and not merely resident,
|
||
and reports the zero-prefill pipe.
|
||
20 is an independent Python re-derivation of the seek-slack
|
||
model, sharing no code with the Lua producer it checks.
|
||
21 decodes the IPL ROM's HD63450 configuration and gates on the
|
||
bytes being where it says they are.
|
||
22 prices a scene change: header bytes, load-time clocks and
|
||
what both cost in accumulated seek slack, across explicit
|
||
rates. Its cycle counts are PARSED out of the rig's log, not
|
||
pasted in, so they cannot go stale silently.
|
||
buscost.py is the shared bus-cycle table. The per-block
|
||
constants live in tools/encoder/vq_hybrid.py and are imported,
|
||
never copied.
|
||
tools/bench/ MAME Lua injection harness and 68000 benchmark sources.
|
||
check.sh is the green light.
|
||
blit.s/blit.lua time the full-frame GVRAM blit on the 68000
|
||
itself. Not part of check.sh, because wall timings would make
|
||
the green light host-sensitive.
|
||
span.sh measures the literal-span mode the same way and
|
||
asserts that every one of its 36 timing configs drew a
|
||
pixel-exact frame, the count taken from generated metadata so
|
||
a new config cannot weaken the gate.
|
||
crtc_mode.lua is the single source of truth for CRTC R00-R08
|
||
and R20. Do not write CRTC values anywhere else.
|
||
prep_dlx.py/decode.lua/verify_decode.py load, time and verify
|
||
decode.s. prep_stream.py/stream.lua do the same for stream.s,
|
||
but lay the container out as a DISK in a host file and feed it
|
||
through a bounded ring at a modelled pipe rate, so the rig is
|
||
not bounded by the emulated machine's RAM and a stock 2 MB
|
||
machine runs the whole window. dlxload.py holds the
|
||
codebook/palette load-time maths both preps share -- and
|
||
the reference src/player/load.i is gated against.
|
||
prep_load.py/load.lua/verify_load.py/load_run.sh run those
|
||
transforms ON the 68000 and compare all 10,752 output bytes
|
||
with dlxload.py's, palette words read back out of the palette
|
||
registers rather than a RAM shadow.
|
||
tools/bench/c68k/ headless px68k C68K harness, a SECOND emulator for every
|
||
68000 cycle figure. Links only px68k's CPU core: no SDL, no
|
||
ROMs, no emulated machine. `make PX68K=~/src/px68k` then
|
||
run.sh; verify_c68k.py checks the decode is pixel-exact, which
|
||
is what licenses the cycle numbers. It also counts BUS cycles,
|
||
which MAME cannot report. The Makefile's -no-pie and the
|
||
harness's MAP_32BIT arena are load-bearing: C68K truncates
|
||
host pointers to 32 bits.
|
||
25 imports nothing itself: it reads the DLXSCENE1 scene
|
||
table and reports the worst gap between two decision points,
|
||
what the input layer has to survive, and what both cost in
|
||
51.3's accumulated slack across explicit rates.
|
||
tools/import/ the ONLY code in this tree coupled to somebody else's source.
|
||
scenegraph.py reads a DirkSimple checkout (and optionally the
|
||
SNES chapter XMLs) and writes tmp/scenegraph.json in this
|
||
project's own DLXSCENE1 schema, with the sources' licences and
|
||
attribution inside it. Nothing is vendored and the output is
|
||
gitignored derived data.
|
||
tools/media/ builds docs/img/ from a paced recording run
|
||
tools/vasm/ vasm m68k assembler, binary plus source tarball
|
||
tools/encoder/ hybrid VQ encoder and DLX3 container writer.
|
||
spans.py is the v7 span geometry, selection and serialiser,
|
||
and the single place the chain layout is stated on the encoder
|
||
side. It must match blit.s and decode.s: 11 coarse units of
|
||
24 px, 11 fine of 2.
|
||
DLX2 4-byte-aligns every frame record, because an odd move.l
|
||
is an ADDRESS ERROR on a 68000, not a slow read. DLX5 aligns
|
||
them to 512 B sectors instead, so a DMA channel can read a
|
||
record as whole sectors straight into the ring with no window
|
||
and no bounce copy; dlx.record_lengths() is the one place that
|
||
rule is applied.
|
||
dlx.py is the reference DECODER, ground truth for the 68000.
|
||
24 models the ring with the 68000 owning it: the request
|
||
queue, the poll-only-when-not-decoding rule and 54.4's frame
|
||
cadence, and reports the pipe the player's own loop gives away.
|
||
src/player/ decode.s is the 68000 DLX3 decoder with a preloaded-stream
|
||
front-end. stream.s is the same decoder behind a bounded ring.
|
||
load.i is the LOAD-time half: codebook expansion and palette
|
||
packing, out of the raw container header, with loadgate.s as
|
||
its rig front-end. Its three scratch tables describe the
|
||
machine rather than the scene, so they are a separate entry
|
||
point a player calls once at boot.
|
||
ring.i is the RING PRODUCER: `aligned` placement, the
|
||
descriptor ring, the prefill policy, 51.2's slack rule as
|
||
arithmetic (ring_may_seek) and a seek. It reads the DLX4 record
|
||
index because a player cannot learn a record's length by
|
||
walking a stream it has not fetched.
|
||
Both include frame.i (the block loop and span chain) and
|
||
geom.i (the constants), so there is exactly ONE copy of the
|
||
bytes every cycle constant is fitted to. The span pass is
|
||
blit.s v7 verbatim, the same instruction sequence the
|
||
66.0/9.143/9.978 clock fit was measured on, so do not tidy it.
|
||
check.sh asserts decode.s still assembles to the same 1,296
|
||
bytes.
|
||
assets/ extracted frames and audio (gitignored)
|
||
```
|
||
|
||
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored. Supply your own.
|