FINDINGS 59.7. tools/analysis/15_bus_occupancy.py has always answered "what does each W cost" and never "what can the frame afford", and after 59.2 those are not the same question. It now answers both, and takes an optional --kbps for the auto-request rows -- the only rows whose cost depends on how long the record takes to arrive. On the gate container at 12 fps, decode term MEASURED: decode 68.5%, audio 1.25%, HEADROOM 30.2% = 6.74 clocks per byte at a 37,403 B record. Against that, P4a cut the transport from 391.7% of a frame to 40..95% -- four to ten times, the largest movement in this project's cost model since the decoder was written -- and it still does not fit. A dual-address byte is a 4-clock read of the device plus a 5-clock write to memory, so 9 clk/B is a FLOOR and the frame affords 6.74. No GCR share goes under it and no delivery rate goes under it: a share decides whether the channel sits at the floor or above it. At 460 KB/s max-rate totals 165.1% and LRAR at 50% totals 117.4%, and a 50% share tops out at 543 KB/s, above which the channel is the bottleneck and the rate falls back to exactly that floor. So 59.2's three bounds arrive in the budget as one sentence: the configurations this machine can run are the ones the frame cannot afford, and the one it can afford -- single address, 5 clk/B, 92.2% total, 7.8% spare -- needs the device to ACK the DMAC, which needs a request line MAME does not connect and the slot pinout does have at B36/B37. ROADMAP re-ranks accordingly. B3 stops being a constant to look up and becomes DOES THE CARD DRIVE #EXREQ, ahead of B1: B1 sets how much headroom the player has, B3 decides whether there is any. New E7 carries the other branch -- if the answer is no, the container must reach 27,995 B a frame, 328 KB/s of payload, against 438 now. The dependency diagram is redrawn around that fork. The scope is stated rather than buried: this is the GATE container, deliberately the heaviest thing the encoder emits, and the lighter cpufit family was NOT priced -- 15_bus_occupancy.py refuses it, correctly, because the C68K measurement it cross-checks against belongs to the gate container. E7 therefore begins with a harness re-run, and until then "34% too big" is a statement about the fixture and not about the project. check.sh ALL GREEN before and after. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
465 lines
28 KiB
Markdown
465 lines
28 KiB
Markdown
# Dragon's Lair: Sharp X68000 port
|
|
|
|
Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SCSI).
|
|
|
|
This is fundamentally a **video codec problem**, not a game-logic problem. The
|
|
game logic is a scene table with branching input windows; the difficulty is
|
|
pushing ~22 minutes of Don Bluth animation through a 10MHz 68000.
|
|
|
|
## What it looks like
|
|
|
|

|
|
|
|
Left, the Blu-ray frame cropped to 256x192. Right, the same frame **as the
|
|
emulated 68000 actually drew it**: 256 colours out of the X68000's 65536, one
|
|
16-colour-per-4x4-block codebook, decoded by `src/player/decode.s` from the
|
|
container. Not a re-render. These are the pixels MAME had on screen, pulled out
|
|
of its own snapshot, 2x nearest-neighbour, no filtering.
|
|
|
|
**The player, running.** 119 frames out of a **256 KB ring buffer on an emulated
|
|
stock 2 MB X68000**, paced to a 12 fps frame clock, streamed from a host file at
|
|
488 KB/s by `src/player/stream.s` with no Lua in the decode path. Source on the
|
|
left, the machine's screen on the right. (This recording was paced by the host;
|
|
the 68000 now keeps that clock itself, off the CRTC's V-DISP, and the same 120
|
|
frames decode pixel-exact under it — `src/player/clock.i`, FINDINGS 54.)
|
|
|
|
<video src="docs/img/player.webm" controls muted loop width="100%"></video>
|
|
|
|
[`docs/img/player.webm`](docs/img/player.webm) (119 frames, 12 fps, VP9)
|
|
|
|
116 of those 119 frames are **pixel-exact** against `tools/encoder/dlx.py`'s
|
|
reference reconstruction. The other three are **torn**: the top of the picture
|
|
is frame *n* and the bottom still holds frame *n-1*, because MAME captured the
|
|
screen while the block loop was partway down it. That is not a rig artefact.
|
|
`decode.s` writes straight to the displayed page, so a real player tears the
|
|
same way. `tools/media/make_readme_media.py` asserts the tear rather than
|
|
trimming it: every differing pixel has to come from the previous frame, or it
|
|
refuses to build.
|
|
|
|
**What the decoder is doing.** The same window with the block-mode map beside
|
|
it. **Black is SKIP** (costs nothing, draws nothing, the previous frame stands),
|
|
**blue is V1** (one codebook index for a whole 4x4 block), **amber is V4** (four
|
|
indices), **red is RAW** (sixteen bytes verbatim). The mode mix is what every
|
|
cost table in `docs/FINDINGS.md` is really about: V4 costs 1.5x V1, and the mode
|
|
decision is charged both bytes *and* cycles, which is why a byte-rich profile
|
|
buys its way out to RAW rather than V4.
|
|
|
|
<video src="docs/img/modes.webm" controls muted loop width="100%"></video>
|
|
|
|
[`docs/img/modes.webm`](docs/img/modes.webm) (the same 119 frames, with the mode map)
|
|
|
|
**Name the layer.** Everything above is **emulated**: MAME 0.277 `x68000`,
|
|
`-bios ipl10`, stock 10 MHz / 2 MB, cross-checked frame for frame on a second
|
|
CPU core (px68k's C68K). Nothing in this project has run on real hardware yet.
|
|
|
|
## Where it stands
|
|
|
|
**The binding resource is the 68000's local BUS, not its clock.** The decoder
|
|
occupies 86.7% of it once instruction prefetch is counted, and 52 of the 53
|
|
frames that miss the 12fps budget miss on the bus (FINDINGS 38). Read that
|
|
before optimising anything for cycles.
|
|
|
|
**The decoder works and is measured.** `decode.s` draws blocks and v7 literal
|
|
spans pixel-exact under both CPU cores, and costs inside the player what the
|
|
standalone blit benchmark said it would, to 0.2% (FINDINGS 41).
|
|
|
|
**The delivery path works too.** `stream.s` decodes the whole 120-frame window
|
|
out of a 256 KB ring on a stock 2 MB machine, final frame pixel-exact, with the
|
|
container in a host file rather than preloaded into RAM. The constraint is
|
|
**contiguity, not byte count**: the block loop reads with a monotonically
|
|
increasing `a0` and no bounds check, so the ring needs the whole next record
|
|
resident *and contiguous*, a condition no byte-counting buffer simulation can
|
|
see (FINDINGS 49).
|
|
|
|
**Seek slack is accumulated, not owned.** A ring's lookahead is built out of
|
|
`pipe - wire` and a seek spends all of it. At 488 KB/s a 256 KB ring needs 4.83
|
|
seconds of play to reach its 7-frame ceiling from empty, and 512 KB needs 8.42
|
|
seconds to reach 14, so a bigger ring raises the ceiling *and* lengthens the
|
|
climb. A branch point therefore asks "has there been enough play since the last
|
|
one", not "is the buffer big enough" (FINDINGS 51).
|
|
|
|
**There is no working delivery rate figure, deliberately.** `--bus`, `--kbps`
|
|
and `DLX_STREAM_KBPS` are required arguments with no defaults, so no table can
|
|
be scored against a rate its own output does not state. What replaces a constant
|
|
is a requirement: `tools/analysis/19_ring_stream.py` reports the **zero-prefill
|
|
pipe**, the rate a medium must clear for a container to need no prefill, which
|
|
is **513.2 KB/s** for the current candidate. That is a hardware acceptance test
|
|
to measure a BlueSCSI against (FINDINGS 50).
|
|
|
|
**The largest open number is W, the clocks stolen per delivered byte.** The
|
|
MB89352 is an 8-bit SPC, so the DMAC pays per byte rather than per word, which
|
|
is a 2x correction the project has already paid for once (FINDINGS 43). What W
|
|
costs is set by how the player programs the DMAC: 5 clocks a byte single
|
|
address with the bus held, 9 dual address held, 12 single address arbitrating
|
|
per byte, 16..19 dual address arbitrating per byte. The design's fate changes
|
|
completely across that ladder, and it is ours to choose.
|
|
|
|
**The one worked example on the machine is expensive.** The X68000 IPL ROM
|
|
programs all four HD63450 channels itself, and
|
|
`tools/analysis/21_iplrom_dmac.py` decodes that configuration out of the ROM
|
|
image and gates on the bytes still being there. Both the audio channel and the
|
|
on-board disk channel are dual address, 8-bit port, cycle steal *without* hold,
|
|
one external request per byte: **16..19 clocks a byte**, the top of the ladder.
|
|
For audio that is a settled figure and a small one, 1.25%..1.48% of a frame. For
|
|
the disk it is where nothing fits at any container size. The ROM drives SASI
|
|
rather than the MB89352, so it does not settle W, but a cheap configuration is
|
|
now the thing that has to be shown rather than assumed (FINDINGS 52).
|
|
|
|
**The player builds its own codebooks and palette now.** The two load-time
|
|
transforms — codebooks to word-per-pixel form, palette to `GGGGGRRRRRBBBBBI`
|
|
with the shared LSB picked per entry — ran host-side until session 21 and now
|
|
run on the 68000, out of the raw container header, byte-exact against the host
|
|
implementation on both CPU cores and with the palette read back out of the
|
|
hardware registers. A scene change costs **18.96 ms**, a third of one 12fps
|
|
frame slot. The finding underneath it is a cost nothing had counted: a scene
|
|
header is **5,920 bytes** that must arrive before frame 0, and in the currency
|
|
of seek slack those bytes lengthen the refill climb by 138 ms at 488 KB/s and by
|
|
**1.099 s at 451.4 KB/s**, because the surplus they are divided by goes to zero
|
|
(FINDINGS 53).
|
|
|
|
**The 68000 fills its own ring now, and the player's request loop costs more
|
|
than the medium does.** `src/player/ring.i` places records, prefills, keeps the
|
|
slack rule and seeks, out of a per-record index the container carries (DLX4).
|
|
The channel only moves bytes while it has a request and only the CPU can issue
|
|
one, so the disc **stands still between records** by an amount the player sets:
|
|
at 488 KB/s a one-deep request queue gives away **6.8% of the pipe and underruns
|
|
59 of 120 frames**, a two-deep one gives away 3.4% and underruns none — on a
|
|
container whose whole surplus over the wire is 8.7% (FINDINGS 55).
|
|
|
|
**The player runs off a real disc now, and PIO costs 87 clocks a byte.**
|
|
`src/player/xfer.i` answers the ring's request mailbox with a real READ(10) to a
|
|
real MB89352 instead of a host moving bytes at a modelled rate: 120 records,
|
|
4,488,588 B, **pixel-exact out of a 256 KB ring**, with a real mid-stream seek in
|
|
a second pass, and the **same 18 wraps** three different transports have now
|
|
produced. What it costs is the finding. Subtracting the same 120 frames run
|
|
twice gives **87.28 clocks per delivered byte**, and the 68000's own cycle table
|
|
for that loop says **87.15** — 0.2% apart, so the cost is the instruction stream
|
|
rather than the emulator's device model, and it is the first number this rig has
|
|
produced that a real board would also pay. At this container's mean record that
|
|
is **391.8% of a 12 fps frame**; the machine's own V-DISP clock agrees from the
|
|
other end at **2.57 fps**. Against the W ladder — 22.4% of a frame at 5 clocks a
|
|
byte, 85.3% at 19 — **the CPU doing the work itself is 4.6x the worst DMA
|
|
configuration this project has found and 17.5x the best.** Getting the DMAC to
|
|
hold the bus is no longer worth 9 against 19; it is worth 87 against either, and
|
|
it is the only thing left before a player (FINDINGS 58).
|
|
|
|
**The DMAC drives the data phase now, and it holds the bus.** `src/player/dma.i`
|
|
programs an HD63450 channel and hands it the SCSI data phase: **the same 2,048
|
|
bytes come off the disc three ways — PIO, the channel with the bus held, the
|
|
channel stealing cycles — and all three are byte-exact.** The evidence that the
|
|
DMAC and not the CPU is driving it never looks at the data register, which
|
|
cannot answer the question: with the DMAC's OWN asserted, MAME cannot tell a
|
|
CPU-driven byte at `$EA0015` from a DMAC-driven one. What it looks at instead is
|
|
**the CPU's own progress**. MTC is sampled by the instruction *after* the one
|
|
that starts the channel; held, it reads **zero of 2,048** — the whole transfer
|
|
happened between two instructions, because the 68000 did not execute in between
|
|
— while the stealing configuration reads the full count and the CPU then goes
|
|
round its own loop 426 times. Put the stealing registers in the held slot and
|
|
the run still delivers every byte and the gate goes **red**, which is what says
|
|
the counter can come out different (FINDINGS 59.1).
|
|
|
|
**And auto-request is charged by time, not by byte.** The card as MAME models it
|
|
has **no request line to the DMAC at all** — its flow control is DTACK — so
|
|
every configuration that can be run against it is auto-request, and an
|
|
auto-requested channel does not know whether the device is ready: it spends its
|
|
share of the bus either way. Every `W` in this project is clocks per *delivered*
|
|
byte, which presumes the device asks; here the cost scales with **how long the
|
|
record takes to arrive**, so halving the delivery rate *doubles* the CPU cost of
|
|
the same record. Priced from the MC68450's own limited-rate constants against an
|
|
explicit 460 KB/s: max rate costs the whole **95.3% of a frame** the record takes
|
|
to land, and of the four bus shares the GCR can be programmed for — 50, 25,
|
|
12.5, 6.25% — **only 50% carries the rate**, at 10.61 clocks a byte and 47.6% of
|
|
a frame. The GCR is a design lever nothing in this tree had named (FINDINGS
|
|
59.3).
|
|
|
|
**And what it all costs: the frame affords 6.74 clocks a byte, and a
|
|
dual-address byte is 9.** Putting the transport on the channel cuts it from
|
|
**391.7% of a 12 fps frame to 40..95%** — four to ten times, the largest
|
|
movement in this project's cost model since the decoder was written — and **it
|
|
still does not fit.** After the measured decode (68.5%) and the audio DMA
|
|
(1.25%), 30.2% of the frame is left, which at this container's 37,403 B record
|
|
is 6.74 clocks a byte; a dual-address byte is a 4-clock read of the device plus
|
|
a 5-clock write to memory, so **9 is a floor no bus share and no delivery rate
|
|
goes under**. Single address is 5 and fits at 92.2% with room to spare — and it
|
|
needs the device to ACK the DMAC, which needs a request line MAME does not
|
|
connect and the slot pinout does have. So the project's live question is now a
|
|
fact about a board: **does a real CZ-6BS1 drive `#EXREQ`?** If it does, the
|
|
design fits. If it does not, the container has to come down from 438 KB/s of
|
|
payload to **328** — which is an encoder target, entirely inside this project,
|
|
and measured against the heaviest container the encoder emits rather than
|
|
against a shipping one (FINDINGS 59.7).
|
|
|
|
**A record is not a sector, and the cheapest fix is a re-encode.** 117 of 120
|
|
records start part way into a 512 B block, and reading whole blocks into the
|
|
ring corrupts the neighbouring records rather than merely wasting bytes — the
|
|
block loop reads with no bounds check. PIO absorbs this for free by simply not
|
|
storing the bytes outside the window, which is a property that disappears the
|
|
moment a DMA channel takes over. Priced three ways: windowed PIO is +1.34% on
|
|
the wire and cannot be done by a channel at all; a bounce buffer is +1.34% and
|
|
**+5 clocks on every delivered byte**, 22.4% of a frame; sector-aligning records
|
|
in the container is **+0.43% and zero clocks**. The last wins on both axes and
|
|
joins the re-encode bundle (FINDINGS 58.3). **Session 27 made it a
|
|
precondition rather than a preference**: the transport now *refuses* a windowed
|
|
read when the data phase is the channel's, so the container has to meet it
|
|
before the DMAC can sit behind the ring (FINDINGS 59.4).
|
|
|
|
**The scene graph is in, and the worst gap between two decision points is
|
|
zero.** `tools/import/scenegraph.py` imports the arcade scene graph — 40 scenes,
|
|
516 sequences, 906 input windows — and 5.4% of the game's 612 branch transitions
|
|
open an input window on the first frame of a clip the disc *seeked to*, so two
|
|
seeks can fall back to back with no play between them. A rule of the form "has
|
|
there been enough play since the last branch" can therefore be answered no by
|
|
the **content**, not by the buffer. It does not break the design: a branch on an
|
|
empty ring costs the 2-record prefill, **149.7 ms at 488 KB/s**, not the climb.
|
|
What it removes is margin — at that rate in a 256 KB ring, **76% of this game's
|
|
branch points arrive before the ring has refilled**, and a 512 KB ring makes it
|
|
90%, because doubling the ceiling does not touch `pipe - wire` (FINDINGS 56).
|
|
|
|
**Nothing outside-derived is committed here.** The scene graph is not
|
|
redistributable from this tree; it is regenerated from a reader's own clones
|
|
into gitignored `tmp/`, and `tools/import/scenegraph.py` is the single file in
|
|
the repo coupled to those projects — everything downstream reads `DLXSCENE1`,
|
|
this project's own schema, with the sources' attribution carried in it.
|
|
DirkSimple is zlib (Ryan C. Gordon); the SNES chapter set is MIT (Chad
|
|
Doebelin) and, by its own README, *derived* from DirkSimple rather than an
|
|
independent transcription, which struck a cross-check this project had planned
|
|
on for eight sessions.
|
|
|
|
**Current encode:** 496.7 KB/s at 29.19 dB, 1 frame of 120 over the 12fps
|
|
budget, and that one is frame 0, the intra frame, late on purpose.
|
|
|
|
**Green-light check:** `./tools/bench/check.sh` (~4 min, needs the Blu-ray
|
|
mounted) re-runs both display regression tests, the rate-control drift gate, the
|
|
display-path coherency counterexample, a 120-frame 68000 decode on two CPU
|
|
cores, the ring and paced-ring passes, the DMAC configuration gate and the
|
|
load-time transforms on both cores, then imports and gates the scene graph
|
|
when a DirkSimple checkout is present, then prints `ALL GREEN`.
|
|
|
|
## Reproducing this
|
|
|
|
**No media ships in this repo and none of it is redistributable.** Bring your
|
|
own Dragon's Lair Blu-ray. Everything else needed to rebuild every number and
|
|
every picture above is either here or is packaged.
|
|
|
|
You need:
|
|
|
|
| | |
|
|
|---|---|
|
|
| the disc | loop-mounted read-only: `udisksctl loop-setup -r -f DRAGONS_LAIR.iso`. The tree was built against a decrypted UDF 2.x image. 7-Zip cannot read UDF 2.x, so use the loop mount |
|
|
| `python3` | plus **numpy** and **Pillow**, and nothing else. The k-means is hand-rolled rather than pulling in sklearn |
|
|
| `ffmpeg` / `ffprobe` | frame extraction, and the clips above |
|
|
| **MAME** | tested on 0.277, with the `x68000` ROM set. The rigs drive it headless via `-autoboot_script` |
|
|
| vasm (m68k, Motorola syntax) | **vendored**: `tools/vasm/vasmm68k_mot` is a Linux x86-64 binary, with the source tarball beside it to rebuild elsewhere |
|
|
|
|
Then:
|
|
|
|
```sh
|
|
export DLX_BDROM=/path/to/your/mounted/bluray # if not /media/$USER/BDROM
|
|
./tools/bench/check.sh # ~3 min, prints ALL GREEN
|
|
```
|
|
|
|
`DLX_BDROM` is honoured by every tool that reads the disc. Two stages are
|
|
optional and **skip rather than fail** when their input is absent, because both
|
|
live outside this repo:
|
|
|
|
- `PX68K=/path/to/px68k` for the second-CPU-core gate. This is the cheapest
|
|
strong test in the tree (seconds, no MAME, no ROMs) and it is what licenses
|
|
the bus and cycle figures.
|
|
- `IPLROM=/path/to/iplrom.dat` for the DMAC configuration gate. Defaults to
|
|
`~/mame/roms/iplrom.dat`.
|
|
|
|
To rebuild the stills and clips in `docs/img/` you also need a paced recording
|
|
run; see the header of `tools/media/make_readme_media.py`.
|
|
|
|
**Scene selection is a hard-coded stream number, not a search.** The gates use
|
|
streams `00020` and `00223` of the disc's 224 `.m2ts` files. A different
|
|
pressing may number them differently, and if so the green light will extract the
|
|
wrong footage rather than fail, so check that `tmp/fr_singe/` looks like the
|
|
Singe encounter before trusting any figure.
|
|
|
|
**Not every large stream is game footage.** `00216` is the feature with a
|
|
burned-in commentary picture-in-picture and `00215` is the commentary itself,
|
|
the two largest files on the disc. The clean 9.4-minute animation is **`00223`**
|
|
(FINDINGS 25.1).
|
|
|
|
## Encoder
|
|
|
|
```
|
|
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
|
|
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile scsi --preview p.png
|
|
```
|
|
|
|
The codec is a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4
|
|
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
|
|
rate-distortion. The RAW escape means `lam=0` is pixel-exact against the
|
|
palettised frame, so the quality knob spans lossless to heavily compressed
|
|
without changing the bitstream.
|
|
|
|
**Two byte budgets, not one.** `--kbps` is the quality rate point and
|
|
`--span-kbps` is the ceiling the span pass may draw on. They are different
|
|
things: the profile is chosen, the pipe is hardware, and bytes between them buy
|
|
a better picture if spent on `lam`, the 68000's deadline if spent on spans, and
|
|
nothing if left unspent. Spans run before `mu` because a span pays in bytes and
|
|
`mu` pays in picture (FINDINGS 41.2).
|
|
|
|
**Two ceilings, on two different axes.** The second is the 68000's decode
|
|
budget: `mu` is bisected per frame against 833,333 cycles so the frame also
|
|
*decodes* in time, which takes the worst sustained window from 37 frames over
|
|
budget to 1, for 0.62 dB at `scsi` (FINDINGS 31). It is on by default and
|
|
`--no-cpu-fit` turns it off. Unlike bytes, cycles have no bucket: there is no
|
|
double buffer to decode ahead into, so it is a hard per-frame ceiling.
|
|
|
|
**One profile, `scsi`, at 280 KB/s.** The 110 KB/s `sasi` profile was dropped on
|
|
capacity rather than bandwidth, since a SASI volume is limited to 40 MB and the
|
|
game's 22.8 minutes is 146 MiB even at that rate (FINDINGS 32). The rate point
|
|
may return under another name once the delivery medium is settled, because a 1x
|
|
CD-ROM sustains ~150 KB/s and CD-ROM is the only period medium with the
|
|
capacity.
|
|
|
|
The profile bitrate is a **ceiling**: `lam` is bisected per frame under a leaky
|
|
bucket, so the profile's `lam` is a quality floor rather than a setting
|
|
(`--fixed-lam` opts out). At `--spans all` none of that binds, though. A
|
|
32-frame bucket emits the same container byte for byte as an 8-frame one and
|
|
`lam` never leaves its floor on any frame of the reference window, because the
|
|
rate is set by the span pass and by `mu` (FINDINGS 44.3). Two known unit
|
|
inconsistencies on that side are implemented and default off because they
|
|
measure as a wash: `--joint-decide` prices a byte at `lam + mu*c` rather than
|
|
`lam`, and `--joint-bucket` stops the bucket lending clocks it cannot repay.
|
|
|
|
An encode is ~95% k-means. A 120-frame window is ~29 s, of which ~22 s is
|
|
training the two codebooks.
|
|
|
|
Profiles are derived from a bandwidth figure rather than chosen by eye:
|
|
|
|
```
|
|
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
|
|
```
|
|
|
|
## Documentation
|
|
|
|
- **`docs/STATUS.md`** is the current state, working setup, blockers and next
|
|
steps. **Start here.** It also lists what has been explicitly abandoned, so
|
|
old ideas do not get re-proposed.
|
|
- **`docs/ROADMAP.md`** is the remaining work to a completion target, and which
|
|
milestone that target is. Read it with STATUS rather than instead of it:
|
|
STATUS holds the measurements, ROADMAP holds the shape and goes stale first.
|
|
- **`docs/FINDINGS.md`** is measured hardware facts, content statistics, the
|
|
codec decision, and a section on measurement traps that produced three
|
|
separate false results. Read §4 before trusting any pipeline number. It is
|
|
append-only and later sections overturn earlier ones; superseded sections
|
|
carry a blockquote pointing at the correction.
|
|
- **`docs/BENCHMARK.md`** is how to measure the storage subsystem, and why a
|
|
bandwidth figure out of MAME would be meaningless.
|
|
- **`docs/HARDWARE.md`** is the X68000 GVRAM/CRTC reference.
|
|
|
|
## Layout
|
|
|
|
```
|
|
docs/ findings, status, roadmap, hardware reference
|
|
docs/img/ the stills and clips above, built from a real emulated run
|
|
tools/analysis/ measurement scripts, numbered in the order they were written.
|
|
Run from the repo root; they import from tools/encoder/.
|
|
01 and 02 are marked BROKEN deliberately and kept as
|
|
regression references.
|
|
10 is a COUNTEREXAMPLE and exits non-zero by design: it
|
|
demonstrates that the two-display-path plan corrupts 70 of 120
|
|
frames, which is why decode.s has one display path.
|
|
15 measures how much of the 68000's local bus the decoder
|
|
occupies and exits non-zero if its derived model stops
|
|
matching the harness's measurement.
|
|
16 is the DLX3 span container round-trip gate: it encodes,
|
|
writes the container, reads it back with the reference decoder
|
|
and fails if a pixel differs, or if it emitted too few spans to
|
|
have tested anything.
|
|
19 models the ring's ADDRESSES rather than its occupancy,
|
|
because each record must be contiguous and not merely resident,
|
|
and reports the zero-prefill pipe.
|
|
20 is an independent Python re-derivation of the seek-slack
|
|
model, sharing no code with the Lua producer it checks.
|
|
21 decodes the IPL ROM's HD63450 configuration and gates on the
|
|
bytes being where it says they are.
|
|
22 prices a scene change: header bytes, load-time clocks and
|
|
what both cost in accumulated seek slack, across explicit
|
|
rates. Its cycle counts are PARSED out of the rig's log, not
|
|
pasted in, so they cannot go stale silently.
|
|
buscost.py is the shared bus-cycle table. The per-block
|
|
constants live in tools/encoder/vq_hybrid.py and are imported,
|
|
never copied.
|
|
tools/bench/ MAME Lua injection harness and 68000 benchmark sources.
|
|
check.sh is the green light.
|
|
blit.s/blit.lua time the full-frame GVRAM blit on the 68000
|
|
itself. Not part of check.sh, because wall timings would make
|
|
the green light host-sensitive.
|
|
span.sh measures the literal-span mode the same way and
|
|
asserts that every one of its 36 timing configs drew a
|
|
pixel-exact frame, the count taken from generated metadata so
|
|
a new config cannot weaken the gate.
|
|
crtc_mode.lua is the single source of truth for CRTC R00-R08
|
|
and R20. Do not write CRTC values anywhere else.
|
|
prep_dlx.py/decode.lua/verify_decode.py load, time and verify
|
|
decode.s. prep_stream.py/stream.lua do the same for stream.s,
|
|
but lay the container out as a DISK in a host file and feed it
|
|
through a bounded ring at a modelled pipe rate, so the rig is
|
|
not bounded by the emulated machine's RAM and a stock 2 MB
|
|
machine runs the whole window. dlxload.py holds the
|
|
codebook/palette load-time maths both preps share -- and
|
|
the reference src/player/load.i is gated against.
|
|
prep_load.py/load.lua/verify_load.py/load_run.sh run those
|
|
transforms ON the 68000 and compare all 10,752 output bytes
|
|
with dlxload.py's, palette words read back out of the palette
|
|
registers rather than a RAM shadow.
|
|
tools/bench/c68k/ headless px68k C68K harness, a SECOND emulator for every
|
|
68000 cycle figure. Links only px68k's CPU core: no SDL, no
|
|
ROMs, no emulated machine. `make PX68K=~/src/px68k` then
|
|
run.sh; verify_c68k.py checks the decode is pixel-exact, which
|
|
is what licenses the cycle numbers. It also counts BUS cycles,
|
|
which MAME cannot report. The Makefile's -no-pie and the
|
|
harness's MAP_32BIT arena are load-bearing: C68K truncates
|
|
host pointers to 32 bits.
|
|
25 imports nothing itself: it reads the DLXSCENE1 scene
|
|
table and reports the worst gap between two decision points,
|
|
what the input layer has to survive, and what both cost in
|
|
51.3's accumulated slack across explicit rates.
|
|
tools/import/ the ONLY code in this tree coupled to somebody else's source.
|
|
scenegraph.py reads a DirkSimple checkout (and optionally the
|
|
SNES chapter XMLs) and writes tmp/scenegraph.json in this
|
|
project's own DLXSCENE1 schema, with the sources' licences and
|
|
attribution inside it. Nothing is vendored and the output is
|
|
gitignored derived data.
|
|
tools/media/ builds docs/img/ from a paced recording run
|
|
tools/vasm/ vasm m68k assembler, binary plus source tarball
|
|
tools/encoder/ hybrid VQ encoder and DLX3 container writer.
|
|
spans.py is the v7 span geometry, selection and serialiser,
|
|
and the single place the chain layout is stated on the encoder
|
|
side. It must match blit.s and decode.s: 11 coarse units of
|
|
24 px, 11 fine of 2.
|
|
DLX2 4-byte-aligns every frame record, because an odd move.l
|
|
is an ADDRESS ERROR on a 68000, not a slow read.
|
|
dlx.py is the reference DECODER, ground truth for the 68000.
|
|
24 models the ring with the 68000 owning it: the request
|
|
queue, the poll-only-when-not-decoding rule and 54.4's frame
|
|
cadence, and reports the pipe the player's own loop gives away.
|
|
src/player/ decode.s is the 68000 DLX3 decoder with a preloaded-stream
|
|
front-end. stream.s is the same decoder behind a bounded ring.
|
|
load.i is the LOAD-time half: codebook expansion and palette
|
|
packing, out of the raw container header, with loadgate.s as
|
|
its rig front-end. Its three scratch tables describe the
|
|
machine rather than the scene, so they are a separate entry
|
|
point a player calls once at boot.
|
|
ring.i is the RING PRODUCER: `aligned` placement, the
|
|
descriptor ring, the prefill policy, 51.2's slack rule as
|
|
arithmetic (ring_may_seek) and a seek. It reads the DLX4 record
|
|
index because a player cannot learn a record's length by
|
|
walking a stream it has not fetched.
|
|
Both include frame.i (the block loop and span chain) and
|
|
geom.i (the constants), so there is exactly ONE copy of the
|
|
bytes every cycle constant is fitted to. The span pass is
|
|
blit.s v7 verbatim, the same instruction sequence the
|
|
66.0/9.143/9.978 clock fit was measured on, so do not tidy it.
|
|
check.sh asserts decode.s still assembles to the same 1,296
|
|
bytes.
|
|
assets/ extracted frames and audio (gitignored)
|
|
```
|
|
|
|
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored. Supply your own.
|