Files
prosolis ba966efe7e Take the player through a branch with sound, and find the predictor does not seek
FINDINGS 71, ROADMAP P6d.  70.3 named exactly what was missing -- packed.s starts
PG_AK/PG_AKF at lump 0 and has no audio seek path -- and priced its absence at a
mean 416.5 ms of silence over the arcade's 409 within-container seek targets.
pg_aseek is that path: the lump index, the stream position, the remainder
accumulator and the byte offset into the group, then the second READ(10) at the
lump's own LBA and a re-arm part way into the buffer.  Measured off a real volume:
132,162 B of spliced stream accounted for byte by byte in MAME's own capture,
across a branch at frame 37 -- four frames into lump 3, deliberately NOT on a
group boundary -- in both chip configurations.  Skip computed 2,604 B, cadence
says 2,604.

THE PREDICTOR DOES NOT SEEK, AND THE ERROR IS DC.  The MSM6258's accumulator is a
pure integrator with no leakage term, so a branch that hands the chip bytes chosen
for a state it is not in produces an offset that does not decay.  Playing through:
DC -355 of 511 with AC 0.00 -- the right shape from the wrong ground -- still -108
four seconds later.  STOP and re-PLAY: all 62,500 post-seek samples are EXACTLY a
decode from the container's own init, and the whole error is the single constant
-65.  A re-PLAY is 5.5x better and neither is zero, so PG_ARST is a mailbox with a
number under it.  The host computes -65 out of the container's bytes and the gate
asserts the equality rather than printing both.

AND THE ONLY FIX THAT REACHES ZERO IS THE ENCODER'S.  A player cannot set the
chip's accumulator, only reset it.  Resetting the encoder's predictor every frame
makes all 119 of the container's branch points exact for 0.33 dB (21.99 -> 21.66),
because the step table's floor is a constant 16.  That is a DLXP3 and it is
deliberately not in tools/encoder.

TWO SILENT BUGS, BOTH CAUGHT BY THE CAPTURE.  pg_udiv32 trashes d4 and pg_aseek
held hz there, so the offset came out 1 byte instead of 2,604 -- 166 ms of the
wrong part of the scene at exactly the right rate, every counter agreeing.  And
one already in the tree that had passed this gate three times: pg_ainit waited on
a READ-BACK MTC before PLAY, which is the same test as "a byte has left RAM" only
if no byte leaves in between.  One does, and the chip then plays the scene one
byte in, forever.  Found by locating the capture's opening samples in the
container image: sector 1 + 1.  The witness is now the count that was written.

ALL GREEN, two new stages included.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-25 15:10:20 -07:00

1036 lines
66 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Roadmap — remaining work to a completion target
Written end of session 19 (2026-08-24), against a tree that is ALL GREEN.
Amended end of session 21: P1 done, P2 half done (FINDINGS 53).
Amended end of session 22: P3 done (FINDINGS 54).
Amended end of session 23: P5 done (FINDINGS 55).
Amended end of session 24: G1 done (FINDINGS 56).
Amended end of session 25: P4 HALF done (FINDINGS 57).
Amended end of session 26: P4b done, P4a is the last open item before M2
(FINDINGS 58).
Amended end of session 27: P4a done at the transport level; THE RE-ENCODE
BUNDLE under P2 is now the only thing between this tree and M2, because 59.4
made sector-aligned records a precondition the transport enforces rather than a
preference (FINDINGS 59). **And 59.7 re-ranks what is left: the frame affords
6.74 clocks a byte, a dual-address byte costs 9, so B3 stopped being a constant
to look up and became the question of whether the design fits at all.**
Amended end of session 28: **THE RE-ENCODE BUNDLE IS DONE — all four items, one
re-measurement (FINDINGS 60).** The container is DLX5, sector-aligned, and the
disc now moves exactly the records. Two of the four closed as NEGATIVES: E2's
`--spans all` default is refused on measurement, and E3's joint span/lam
selection emits byte-identical containers because `lam` never leaves its floor.
**What is left of M2 is P4a's wiring** — the DMA channel behind `ring.i`'s
mailbox — and the budget did not move: headroom 6.74 -> 6.69 clk/B, so every
conclusion in 59.7 stands.
Amended end of session 29: **THE DECODER-FREE PACKED PLAYER IS BACK, MEASURED,
AND IT FITS THE CLOCK BUDGET THE CODEC MISSES (FINDINGS 61).** The packed
full-frame blit is **27.3%** of a 12 fps frame — measured, not assumed — a
channel fills GVRAM in buffer mode off the disc with the CPU halted, and it
walks the 1,024 B line stride itself through array chaining. At the 9 clk/B
floor the codec is 110.4% and a decoder-free packed player is **55.2%**. It asks
**576 KB/s, fixed**, against E7's 327 KB/s target. **So B1 stopped setting how
much headroom the player has and started deciding WHICH PLAYER EXISTS**, and B2
stopped being a nice-to-have. The codec cannot be packed: 47.6.4 is closed and
the answer is no, both ways (61.3).
Amended end of session 30: **K1 IS DONE AND THE ANSWER IS THE GOOD ONE
(FINDINGS 62).** A channel writes the palette registers at `$E82000` byte-exact,
and ONE array-chained start crosses from device registers into GVRAM — so a
frame is a palette entry and 192 row entries, started once, with the CPU halted
throughout, and the array is scene-constant. **It opened B4**: MAME models the
palette as plain `COMBINE_DATA` storage with no handler to be wrong about, so
the run bounds the model and not the board, and what a real palette register
does with a byte write is UNMEASURED. B4 is the cheapest hardware item in the
project and a negative costs 0.28% of a frame. **K2, the packed container, is
next.**
Amended end of session 31: **K2 IS DONE, AND IT COST THE BRANCH SOMETHING
(FINDINGS 63).** DLXP1 is a 49,664 B record = 97 sectors exactly, no index, no
decoder, 582.0 KB/s — exactly 61.9's prediction — and px68k's own `gvram.c`
renders the container's bytes index-exact with the harness computing no
interleave. 61.9's picture claim survives the real builder: **34.05 dB against
its 34.08**, and the hardware GRB555 word is charged on top for the first time
in this project (0.53 dB, on every row, so it moves nothing). **But the control
row landed exactly on the codec's ceiling**, so the whole +2.31 dB is the
PER-FRAME PALETTE and nothing else — and 90% of that palette changes every
frame, which makes a mismatched paint **12.8 dB worse** than the correct
pairing. **B2 now decides which packed CONTAINER exists, not only which player**
(63.4). The fallback is a flag: `--scene-palette --no-palette` is 30.79 dB, zero
churn, 576.0 KB/s and still +2.07 dB on the shipping codec. **K3 is next.**
Amended end of session 32: **K3 IS DONE, AND IT FOUND THE THING THAT DECIDES
THE BRANCH (FINDINGS 64).** `src/player/packed.s` brings up its own display,
builds its own 193-entry chain, keeps its own clock off V-DISP and fetches every
record itself off a CZ-6BS1: **120 of 120 frames pixel-exact, every one
compared, in both palette orders.** Two things came with it. **(1) The write
window is the frame.** A packed write needs R20 bit 11, buffer mode blanks the
layer, and a DMAC-direct player holds the window open for the whole data phase
— so the dark fraction of a slot is `record / (DATA-PHASE rate x slot)` and a
medium that exactly meets the 582.0 KB/s sustained requirement **displays none
of the frames it delivers on time**. The rate that matters here is the BURST
rate, which is a third hardware number **B1 has no test for**. It also
**reverses 61.5's ranking**: a packed player that DMAs to RAM and paints with
the CPU opens the window only for the measured 27.3% blit, so it is on screen
72.7% of every slot at any rate, and the two are equal only at **2,131 KB/s =
3.7x the wire**. **(2) A held channel costs the frame clock half its ticks and
the clock cannot tell** — 46.9% of V-DISP edges lost, zero late frames reported,
the player believing 12 fps while the screen ran at 6.37. **The open item is now
K4.**
Amended end of session 33: **P6 HAS AN ENCODER AND A PRICE (FINDINGS 65).**
`tools/encoder/adpcm.py` is an MSM6258 codec whose in-loop decoder is gated
SAMPLE-EXACT against ffmpeg's `adpcm_ima_oki` — there is no reference encoder
for this format, so that is the only check available and it is the one that
catches an encoder agreeing with its own wrong decoder. The Singe window encodes
to **78,125 B at 21.97 dB**, which is 7,812.5 B/s to the byte and 52's figure
arriving from the other direction, and **normalising the disc's 13.4 dBFS level
buys 0.00 dB**, so the level is not a lever. Two things came with it. **(1) The
two published delta formulas disagree by at most 3 in 12-bit units and that is
worth 25 dB** — encode with one and decode with the other and the SNR goes from
21.97 dB to **2.88 dB**, because ADPCM is recursive the way the codec is
temporally recursive (64.1). Which one the chip runs is now **P6a** and it is a
precondition on shipping any audio. **(2) The packed container's own best
property is what makes audio cost.** A record has no index BY DESIGN, so audio
cannot be per-record without making records variable; it rides a fixed cadence
`(F, A)`, the obvious cadence F=1 wastes **57.3%** of every audio sector, and the
pick is **F=11, A=14 — 0.09% padding, 14,336 B held, wire 582.0 → 589.6 KB/s**.
The codec container, which already has an index and variable records, pays
**zero** padding. That is the first cost anyone has found for the packed
branch's simplification, and 64's risk list predicted there would be one without
knowing what.
Amended end of session 34: **P6a IS DONE, AND THE ENCODER WAS WRONG ON FOUR AXES
RATHER THAN ONE (FINDINGS 66).** 68000 code programs HD63450 channel 3 with the
IPL ROM's own ADPCM bytes — dual address, 8-bit port, cycle steal, EXTERNAL
request — feeds the MSM6258 at the chip's own pace (**7,811.4 B/s against
7,812.5**), and one of **sixteen** candidate decoder models reproduces MAME's
capture **sample-exact over 1,678 consecutive samples**, with a negative control
on every axis. The chip runs **`terms`, LOW nibble first, a 10-bit accumulator,
starting at 2**, and `tools/encoder/adpcm.py` defaulted to the opposite of all
four. **65.2 named the wrong axis as the risk**: the delta formula is worth
2.88 dB and the NIBBLE ORDER is worth **25.74 dB**, and 65.1's "high first,
measured" was a measurement of ffmpeg's VOX file convention rather than of a
chip's data register. Two things came with it. **(1) The 10-bit clamp is free on
the Singe window and only because that window peaks at 435 of 511** — 1.4 dB of
headroom on a 13.4 dBFS passage, where the encoder had been clamping 12.1 dB
higher — so **the audio LEVEL is an open choice again**, downward, and the
loudest passage on the disc is unmeasured. **(2) The transport is P6b's, not
scaffolding**, and it worked first time. **P6b is next and its bytes are
decided: DLXP2 encodes with `adpcm.CHIP`.**
Amended end of session 35: **P6b IS DONE AND ITS PADDING WAS DRIFT (FINDINGS
67).** DLXP2 exists: a 64-byte header, groups of one `A`-sector audio lump then
`F` records, `record i = off_frm + i*rec + (i//F)*A*512` — **still no index and
still none needed**, which is the packed branch's whole claim surviving the one
change that could have ended it. The player carries it in **six instructions
once a frame**, and 120 of 120 records are still pixel-exact off a real volume
with the interleave in, against a **silent control** that says no picture byte
moved. **65.3's 0.09% was not waste, it was a rate error**: a lump is 7,168 B of
space and eleven frames of audio is 7,161.4583… B, so a player that fed the chip
the whole lump would run 0.09% fast — **1.25 s of lip-sync over the game's 22.8
minutes**. The payload is a remainder, `acc += 11*15625; n = acc//24; acc %= 24`,
which is `clock.i`'s shape for `clock.i`'s reason and the **third** time this
tree has met the pattern. The four ADPCM axes ride in the header and the gate
flips each one to prove they earn it (**order 31.99 dB, formula 24.86**). The
wire is **589.6 KB/s**, 65.3's figure to the tenth. ~~**What is left of P6 is the
last quarter: no audio has been played out of this container on any layer, and
two DMA channels have never run at once.**~~
Amended end of session 36: **P6c IS DONE, AND HOLDING THE BUS COSTS THE AUDIO
463x THE SEAM (FINDINGS 68).** All 78,125 B of the container's audio reached the
chip, in order, sample-exact against the four axes in its own header, out of a
player that was fetching records off the same disc at the same time. **Two DMA
channels have now run at once and the interaction is not in the bytes, it is in
the timing**: the MSM6258 has no starvation state, so the interval between a
channel counting out and the CPU arming the next lump is a held nibble pair
driving the predictor. Stealing, the seam is **0.51 ms over ten seconds**;
**held, it is 236 ms — 2.31% of the audio, every one of the ten lump boundaries,
worst 72.8 ms**, because a burst channel halts the 68000 and the only place a
68000 driving this video path has time is inside the transfer wait (250,000 of
250,240 service calls). That is FINDINGS 64.3 reaching the audio: **the held
configuration cannot feed a second consumer**, and B1's answer now decides a
sound as well as a picture. 67.2's drift is measured rather than derived —
predicted 1.25 s, played **1.26 s**. **What is left of P6 is the level (66.3)
and the refill climb with a second consumer through a real branch point.**
Amended end of session 37: **THE AUDIO LEVEL IS MEASURED, AND THE ANSWER IS
THAT IT DOES NOT CHANGE (FINDINGS 69).** Every stream of the game's own footage
(`00000`-`00201`) through `extract_audio.py`'s chain says the disc peaks at
**946 of 2048 = 6.71 dBFS**, which is **5.35 dB over** the chip's 10-bit clamp
— and the census behind that peak is **687 samples of 20.2 million, 402 events,
44.0 ms in 21.5 minutes**. Forty windows drawn over the game and encoded at six
gains then price the choice: the disc's own level has the best mean SNR
(**22.03 dB**) and the gain that guarantees zero clamping costs **0.85 dB across
the whole game** to buy back **1.90 dB on the 2.11 s that clamp**, because the
OKI step floor is a constant 16 and does not scale with the signal. **And
66.3's mechanism does not survive a control**: error after a clamp run is
elevated ~5x, and so is the same window at a gain that never clamps, read at the
same indices — worst ratio **1.28** — because `adpcm.encode` runs the chip's
clamp inside its own search and therefore never loses the chip's state. The
worry was right about the mechanism and aimed one layer too late; session 34
had already closed it. `--audio-gain` exists so the level is a parameter with a
measurement behind it, its default is 1.0, and the gate container is
byte-identical. **What is left of P6 is the refill climb with a second consumer
through a real branch point** — and 69.4 adds two small unbuilt cases, a scene
with no audio track (`00176`) and a scene whose audio is shorter than its video
(`00199`).
Amended end of session 38: **P6 IS CLOSED, AND THE LAST ITEM IN IT REOPENED THE
CADENCE (FINDINGS 70).** `tools/analysis/36_branch_audio.py` runs 51.3's climb
against 56.3's 612 real branch points with the second consumer on the wire. On
the codec branch **audio is 1.7% of the wire and up to 3.30x of the climb**
6.70 s → 22.11 s at 451.4 KB/s, and 79% → 99% of the game's branch points
arriving under it — because slack accrues out of `pipe wire` and that is a
small difference of two large numbers. On the packed branch **there is no climb
at all**: a record goes straight into GVRAM, video lookahead is zero records,
and the only consumer with any lookahead is the audio one (1.833 s of sound
against 0.000 s of picture). **At 589.6 KB/s — the acceptance figure — the
sounded packed container's surplus is exactly zero**, which is the sharpest
statement yet of why B1 must name a burst rate and not only a sustained one.
And a cost nobody had counted: a DLXP2 group is `lump k, then F records`, so a
seek into a group finds its audio BEHIND it — **mean 416.5 ms of silence at the
409 within-container branch points, worst 833.3 ms, 36 of 409 free**, while all
203 scene changes are free by construction. **The fix is one extra read of
7,168 B (11.7 ms against 416) and no player has an audio seek path**; the
alternative is F=1, which gives back 12,288 B of RAM for +4.36 KB/s of wire.
**The cadence pick is therefore reopened and the deciding number is B1's**: what
one extra SCSI command costs. **What is left of P6 is nothing** — 69.4's two
unbuilt cases and 70.3's seek path are player work, not open questions.
Amended end of session 39: **P6d IS DONE — THE PLAYER SEEKS WITH SOUND ON IT —
AND WHAT IT UNCOVERED IS AN ENCODER ITEM (FINDINGS 71).** `pg_aseek` rebuilds
the lump index, the stream position, the remainder accumulator and the byte
offset into the group, issues the second READ(10) 70.3 asked for, and re-arms
channel 3 part way into the buffer: **132,162 B of spliced stream accounted for
byte by byte in MAME's own capture**, across a branch at frame 37 that is
deliberately *not* on a group boundary, in both chip configurations. 70.3's
416.5 ms of silence is gone. **What replaces it is smaller and is not a player
problem: the MSM6258's accumulator is an integrator with no leak, so a branch is
a DC offset that does not decay** — 355 of 511 playing through (still 108 four
seconds later, AC 0.00), an exact permanent 65 if the chip is STOPped and
re-PLAYed. **A re-PLAY is 5.5x better and neither is zero.** The only fix that
reaches zero is the ENCODER's: reset the predictor where a branch can land, and
**resetting every frame makes all 119 of the container's branch points exact for
0.33 dB**. That is a DLXP3 and it is now the cheapest open item in the tree. The
seek also found a race that had passed this gate three times — `pg_ainit` waited
on a *read-back* MTC and could start the scene one byte in, silently — which is
71.2 and is fixed.
**THE COMPLETION TARGET IS M3, THE VERTICAL SLICE** (USER DECISION): one scene
tree — a decision point, two outcomes, a death clip — with audio, streaming from
a real SCSI volume on a stock 2 MB machine, playable. That is the point at which
every layer of this design has been shown to work at once. M4 is listed because
it is real work, but past M3 it is content grinding rather than open questions.
`docs/STATUS.md` remains the session-by-session record and the handoff. This file
is the shape of what is left; where the two disagree about what is done, STATUS
is the one with the measurements and this one is the one that goes stale. Both
were wrong about two encoder gaps until this file was written — see "What was
already done" below.
---
## Status of the four resources
The project's own framing, restated because every item below is priced in one of
these units:
| resource | state |
|---|---|
| **68000 local bus** | the binding one. Decoder occupies 86.7%; 52 of 53 missed frames miss on the bus, not the clock (FINDINGS 38). |
| **68000 clocks** | measured, and the rate controller binds on them. |
| **Delivery rate** | **no working figure, deliberately** (FINDINGS 50, USER DECISION). Every tool REQUIRES an explicit rate. |
| **Seek time** | **no figure at all, and never had one.** 51.3/51.4 made it matter. |
| **W, clocks stolen per delivered byte** | 5 single-address held, 9 dual held, 12 single arbitrated; the IPL ROM's own disk channel is **16..19** (52.5). **The largest open number in the project.** Session 27 added the row underneath it: with **no external request line** on the card (59.2) the channel is auto-requested and is charged **by time rather than by byte**, so at 460 KB/s a 50% bus share costs **10.61 clk/B** and a smaller share cannot carry the rate at all (59.3). |
| **The frame's headroom for a transport** | **6.69 clk/B** — 30.2% of a 12 fps frame, after the MEASURED decode (68.6%) and best-case audio (1.25%), at the DLX5 gate container's **37,585 B delivered record** (session 28: the sector pad is delivered, so `15_bus_occupancy.py` charges it). **It is the number every row above is read against**, and a dual-address byte's floor is 9. It was 6.74 against the DLX4 container; the bundle moved it by 0.05 and moved no conclusion. |
---
## What was already done, and was still on the list
Found while inventorying for this file. Both had been closed in code for several
sessions and were still listed as open gaps in `docs/STATUS.md`:
- **4-byte record padding.** `DLX2`, `encode.py:139-156`, inside rate-control
accounting, reported per frame and per second.
- **CPU cost in the mode decision.** `vq_hybrid.py:218`, priced against measured
per-mode cycles with the exact clustered SKIP rule.
Both entries are now struck in STATUS. **The lesson is procedural: a gap list
that is only ever appended to manufactures phantom work.** Anything crossed off
below should be crossed off in STATUS in the same sitting.
---
## Blocked on hardware this tree does not have
None of these block M2 or M3 software work, because session 18 forced every rate
to be an explicit argument. They set constants, and two of them decide how much
headroom the finished player has.
**B1. Measure the BlueSCSI — throughput, seek time AND the DATA-PHASE BURST
RATE.** The third one is session 32's (FINDINGS 64.2) and it is not a refinement
of the first: sustained throughput decides whether record *i* arrives before
slot *i*, and the **burst rate during the data phase** decides how much of the
slot the picture is on screen, because a DMAC-direct packed player holds the
GVRAM write window open for exactly as long as the transfer takes and buffer
mode blanks the layer. A drive with a read-ahead cache can pass the first and
fail the second. The acceptance test is **`record / (burst x slot)` = the dark
fraction**; at the container's own 582.0 KB/s it is 1.0, and the picture is never
displayed. **Session 29 gave this
a second acceptance test that is not a codec figure at all: 576 KB/s SUSTAINED,
which is what a decoder-free packed literal frame costs and cannot be talked down
from (FINDINGS 61.5).** A codec's bitrate is a lever; a literal frame's is
geometry. So the measurement now has three thresholds to be read against —
453.6 KB/s (the gate container needs no prefill), 327 KB/s (E7's target at the
dual-address floor) and 576 KB/s (no decoder at all) — and which of them the
medium clears decides which player gets built.
Throughput has an acceptance test already derived from real record sizes:
**513.2 KB/s** for the session-14 candidate, **453.6 KB/s** for the DLX5 gate
container (`19_ring_stream.py`, FINDINGS 49.5; the figure was 451.4 before
session 28's re-encode and the sector pad raised it). Seek time has nothing.
51.3/51.4 is why the second half matters: slack is *accumulated* out of
`pipe - wire`, so what a branch point costs is set by the rate and the time since
the last branch, not by the ring size. At 460 KB/s every ring from 192 KB to
512 KB is rate-bound and never fills. **Do not substitute a guess** — run at
several explicit rates and report the sensitivity. That is exactly how the
retired pipe constant survived five sessions after 42.1 called it folklore.
**B2. Does buffer mode blank the display?** `probe_bit11_blank.lua` is written
and settles it in minutes on a real board. FINDINGS 48 shifted the prior toward
MAME and toward "unusable" — **do not pre-build on 1.0 B/pixel**. Same sitting:
the priority register `0xE82500` at `0x0000` (47.3).
> **Session 29 raised what this is worth, and gave it a number to be worth
> (FINDINGS 61).** It used to gate a derived halving. It now gates a player that
> has been measured to fit a budget the shipping design misses — 55.2% of a frame
> against 110.4% at the 9 clk/B floor. And 61.6 found an asymmetry worth carrying
> to the board: the black interval is the PAINT, not the frame, so a CPU-painted
> packed player is dark for **27.3%** of a frame while the cheaper DMAC-direct one
> is dark for **30..113%**. Under MAME's reading the cheap architecture is the
> dark one. **B2 and B1 are now the same decision from two sides, and B2 is the
> five-minute half.**
>
> **Session 31 raised it again, from the other direction (FINDINGS 63.4).** If
> buffer mode does NOT blank, the packed player's per-frame palette is not
> merely visible during the paint — **231 of 256 entries change every frame**,
> and a picture under the neighbouring frame's palette is **12.8 dB worse** than
> the correct pairing, on screen for roughly half of every frame slot at 12 fps.
> So a non-blanking board does not just cost the packed player a blank interval,
> it may cost it the per-frame palette — which is **the whole +2.31 dB** the
> branch has over the codec's ceiling. The fallback is already an encoder flag
> (`pack.py --scene-palette --no-palette`: 30.79 dB, no churn, 576.0 KB/s), so
> what B2 decides is now **which packed container ships**, not whether one can.
**B4. Does a real palette register take a BYTE write?** (62.4, new session 30.)
`$E82000` is 256 16-bit registers. The decoder-free packed player's per-frame
palette rides the frame's array chain as a 193rd entry, which means a
dual-address channel with an 8-bit device port writes it **one byte at a time**,
even bytes into the high half of a register and odd into the low. **MAME cannot
be asked**: it maps the palette to `palette_device::read16/write16`, which is
`memory_array`'s `COMBINE_DATA` over plain storage — RAM that honours
`mem_mask`, with no handler that could refuse. So the run in 62 is a statement
about the model and not about the board.
**It is the cheapest item on this list, cheaper than B2**: write `$A5` to
`$E82000` and `$5A` to `$E82001` from the CPU and read the word back. If it
comes back `$A55A` the palette rides the chain. **The blast radius of a negative
is small and known** — the palette leaves the chain and the CPU writes 256 words
a frame at 61.9's derived 0.28% of a frame — which is why this is B4 and not
above B2.
**B3. Single-address vs dual-address DMA — and now, first, DOES THE CARD DRIVE
`#EXREQ`?** 242 KB/s and 0.69 dB was the old framing. **Session 27 promoted this
to the item that decides whether the design fits at all (59.7).** The frame
affords **6.74 clocks a byte**; a dual-address byte is **9** — a 4-clock read of
the device plus a 5-clock write to memory — so **no dual-address configuration
fits this container at 12 fps, at any delivery rate and any GCR share.** Single
address is 5 clk/B and fits at 92.2% of the frame with 7.8% to spare, and it
needs the device to ACK the DMAC directly, which needs the request line.
**MAME cannot answer it**: `x68k_scsiext.cpp`'s `drq_w` only stores a flag and
the expansion slot has no request path to the HD63450 at all (59.2). The slot
PINOUT has `#EXREQ` at B36 and `#EXACK` at B37, so the provision exists on the
real board. **What is wanted from hardware is therefore narrower and sharper
than "disassemble the ROM": does a CZ-6BS1 assert `#EXREQ` during a data phase,
and will the HD63450 run a single-address channel against it?** A scope or a
logic analyser answers the first; the second is a program the player can run.
Sourcing `scsiexrom.bin` (8 KB, CRC `7be488de`) and disassembling its DMAC init
is still the cheapest paper route to the same answer, because whatever Sharp's
own driver programs into DCR's DTYP is a statement about what the card supports.
FINDINGS 48.4. Not on this machine (checked, session 18).
**Ranking, amended session 29.** It was: B1 sets how much headroom the player
has, B3 decides whether there is any. FINDINGS 61 adds a third reading — **B1
and B2 together decide which player exists.** If the medium clears 576 KB/s
sustained and buffer mode does not blank, the decoder-free packed literal fits at
the dual-address floor B3 cannot get under, and B3 stops mattering for video at
all. If it does not, B3 is still the question. The three hardware facts are no longer
independent, and B2 is by far the cheapest of them.
> **ENCODER WORK IS PARKED — USER DECISION, session 29.** The first draft of this
> amendment said "nothing here is a reason to stop work on the codec". That does
> not survive its own arithmetic. It rested on comparing the packed player's
> 576 KB/s against **E7's 327 KB/s target, which does not exist**: the codec that
> exists is 440 KB/s and 110.4% of a frame, so the real gap is 1.31x, not 1.76x.
> And the branches are not symmetric. **Packed needs two facts** — buffer mode
> does not blank, medium clears 576 KB/s. **The codec needs E7 to succeed** —
> unproven, and 60.4/60.5 measured two of its three levers inert — **AND** the
> medium to clear 327, **AND** it ships at ~100% of the frame with no margin,
> which is where 55.2% is now.
>
> **E7 and E4 are both parked**, E4 included: `H.build`'s k-means builds VQ
> codebooks and a literal player has no VQ. C1 is gated by E4 and follows.
>
> **The codec is KEPT AND NOT BUILT ON.** That is inventory, not work. B2 is
> unanswered and 48.1's prior leans against packing — an assertion against a
> silence — and if buffer mode blanks there is no version of the packed player
> that is merely expensive (48.3), at which point the codec is the only path
> left. Keeping a working decoder on disk costs nothing; building on it costs
> sessions.
**The older ranking, which still holds inside the codec branch:** B1 (throughput
and seek) sets how much headroom the finished player has. **B3 decides whether
there is any.** If the card drives `#EXREQ`, the ladder applies and the design fits with
room. If it does not, the fallback is limited-rate auto-request at a share the
player picks (P4c), and the container has to come down to **328 KB/s of payload**
to fit at the 9 clk/B floor — 34% below where the gate container sits (59.7).
> **Session 20 moved the prior hard, and it moved the wrong way (FINDINGS 52.5).**
> The IPL ROM *is* on this machine, and `tools/analysis/21_iplrom_dmac.py` reads
> its HD63450 setup: the on-board disk channel (ch1, SASI) is `DCR = $80` —
> **dual address, 8-bit port, cycle steal WITHOUT hold**, with `REQG = 10`
> external request, i.e. a full bus arbitration per byte. That is **16..19
> clocks per delivered byte**, above the whole 5..12 bracket 42.4 costs P4 in.
> Same vendor, same DMAC, same class of 8-bit port — but it is *not*
> `scsiexrom.bin`, so B3 stays open. What it changes is that a cheap
> configuration is now the thing that has to be **shown**, not assumed.
---
## The packed branch — what building it means (session 29, USER DECISION)
**This is where the work goes now.** FINDINGS 61: a decoder-free packed literal
player is **55.2% of a frame at the 9 clk/B dual-address floor** against the
codec's 110.4%, and **+4.89 dB** on the shipping container because a literal
frame is not tied to a scene palette the codec's codewords index into. It costs
**582 KB/s, fixed, with no lever.** Three items, in order.
~~**K1. Can a DMA channel write the palette registers at `$E82000`?**~~
**DONE, session 30 — FINDINGS 62. YES, in this model.** `dmagate.s` runs 79:
512 B off the disc into the whole graphic palette, **byte-exact in 256
register words** read back out of `$E82000` by the 68000; the same transfer
aimed at RAM leaving the palette as the CPU poisoned it, which is what
attributes the first run to the channel's `MAR`; and **ONE array-chained start
crossing from device registers into GVRAM**, which is the shape of a whole
frame. The destination is POISONED first (62.1) because "it matches" was a
weak claim against a record that is mostly pad, and the host counts the
poison's discriminating power rather than assuming it: 511 of 512.
**And the array is SCENE-constant** (62.3) — the packed layout spends both
256-colour pages, so there is no page to flip and the 193 destinations never
change; the 1,158 B array is built once at scene setup. What is left on the
CPU per frame in the video path is the channel start and the READ(10), and
neither is priced — say "no per-frame PAINT work", not "no per-frame CPU work".
**It opened B4** (62.4): MAME models the palette as a generic `palette_device`
over `memory_array`, whose `write16` is a plain `COMBINE_DATA`, so it has no
handler to be wrong about and cannot discriminate. And it filed one open
design choice, 62.5: palette FIRST or 193rd is visible on screen and is not
decided. **Session 31 PRICED it and it is a wash — 20.32 dB against 20.33
(63.4) — so it is a container flag (`--palette-last`) and K3 runs both.**
~~**K2. A packed container.**~~
**DONE, session 31 — FINDINGS 63.** `tools/encoder/dlxp.py` is DLXP1 and
`pack.py` writes it: 254 colours with index 0 held free and black at 255,
column *i* interleaved with *i+128*, a per-frame palette, and a **49,664 B
record that is 97 sectors EXACTLY** — the alignment that cost session 28 a
re-encode is free here because a packed record's length is geometry. **No
index and no length word**, for the same reason: nothing has to be walked.
**582.0 KB/s, which is 61.9's prediction to the tenth.** Encodes in 3.3 s
because there is no k-means in it.
**The re-derivation is done and 61.9 survives it: 34.05 dB against 34.08**,
and the GRB555 word is charged for the first time (63.3) — 0.53 dB, on every
row of the table, so it moves no comparison. **The two reserved entries cost
0.0003 dB.**
**What it also found** (63.4): the SCENE-palette control lands exactly on the
codec's ceiling, so the whole +2.31 dB is the per-frame palette; 90% of that
palette changes every frame; and a mismatched paint is 12.8 dB worse, for
roughly half of every frame slot, if buffer mode does not blank.
- ~~**K3. End to end, off the disc.**~~ **DONE, session 32 — FINDINGS 64.**
`src/player/packed.s`, `tools/bench/packed.lua`, `tools/bench/packed_run.sh`,
`tools/bench/verify_packed.py`, `tools/analysis/31_display_duty.py`. Palette,
page-1 X-scroll 384, priority `vc1 = 0x0002`, R20 bit 11, one chained DMA a
frame, **120 of 120 pixel-exact in both palette orders** — and the gate checks
every frame rather than the last, because a packed frame is a literal and the
codec's recursion was what made one comparison audit 120.
**What it found is K4's whole content**, and it is in the amendment above: the
window is the frame, the burst rate decides visibility, and a held channel
eats the clock.
- **K4. THE PACKED PLAYER THAT IS ON SCREEN.** 64.2's option B: DMA the record
into one of two RAM buffers with the window SHUT, then paint it with the
packed `movem` blit (`blit.s` V8, **measured** at 227,553 clocks = 27.3% of a
slot). **82.2% of a frame at the 9 clk/B dual-address floor** against A's
54.9%, **99,328 B of RAM**, and a picture on screen **72.7% of every slot at
any delivery rate** instead of 0% at the container's own wire.
It is not a rewrite of K3: `packed.s` keeps its display bring-up, its clock,
its transport and its record arithmetic, and what changes is the chain's
destination and the addition of a paint. **The one thing in it that has never
been run is the overlap** — a channel filling buffer *i+1* while the CPU
paints buffer *i*, which is the first time in this project that the DMAC and
the 68000 have had to want the bus at the same time for a whole scene.
**K4 is conditional on B2 the same way K3's ranking was**, and the condition
now cuts the other way: if buffer mode does NOT blank, A is on screen the whole
slot and K4 is 27.3% of a frame spent on nothing. **Do not build K4 before B2
is answered** — that is the same rule 61.7 wrote for the codec, applied to the
branch that replaced it.
**What K3 deletes, and why that is a risk and not a win to be banked:** a
DMAC-direct packed player has **no ring**`ring.i`, `xfer.i` and most of
`stream.s` leave the video path, and **P4a's wiring is parked with them.** A
simplification that large usually hides something, and 61.7.2 names the specific
untested thing: a chained transfer has never run back to back at 12 fps.
**K1 and K2 survive a bad answer to B2. K3 does not.** ~~Do K1 first.~~ ~~K1 is
done (session 30, FINDINGS 62); K2 is next.~~ **Both are done. K3 is next — and
63.4 added a second thing for it to run: BOTH palette orders, which is a flag
(`--palette-last`) and not a re-encode.**
---
## M2 — a player, as opposed to a decoder
`decode.s` draws pixel-exact frames from RAM Lua pre-loaded; `stream.s` decodes
out of a bounded ring fed by a host file on a paced clock. Neither is a player.
**Exit criterion: boots from a real SCSI volume on a stock 2 MB X68000, plays
one scene at 12 fps from disc, no host-file pipe, no Lua in the loop. Silent.**
~~**P1. Codebook expansion on the 68000.**~~ **DONE, session 21 — FINDINGS 53.**
`src/player/load.i` expands both codebooks out of the raw container header,
byte-exact against `dlxload.py` on both CPU cores. **9.26 ms**, and it was
priced where it lands rather than treated as free setup: the scene header is
**5,920 B that no rate table in this tree counted**, and in the currency of
51.3 — accumulated slack — those bytes lengthen the refill climb by 138 ms at
488 KB/s and by **1.099 s at 451.4 KB/s**, because the surplus they are divided
by goes to zero. The whole fixed cost of a scene change is about a third of one
frame slot; what makes a branch point expensive is still the seek and the climb.
Shipping the codebooks pre-expanded was considered and refused: it trades
9.26 ms of CPU for 5,120 more header bytes, which is a wash in milliseconds and
not a wash in kind (53.6).
~~**P2. Palette packing on the 68000.**~~ **DONE, session 28 — FINDINGS 60.**
The half that was open was the encoder's, and it closed with the whole bundle
below. Session 21 — FINDINGS 53 — did the packing itself:
~~The encoder still emits RGB888; the X68000 word packing is Lua-side.~~ The
packing is on the 68000: `pal_pack` writes 256 words straight into `$E82000`
with `I` chosen per entry by minimum squared error (**1.96 dB**, 23.3), gated on
the words read back **out of the palette registers**. 9.70 ms per scene, plus
5.29 ms of scene-independent table build hoisted to boot (53.3).
~~**What is left is the other half of the sentence: reserve index 0 as black with
`I = 0` (23.4).**~~ **DONE, session 28 — 60.3.** `VQ.scene_palette` quantises
the picture into 255 entries and holds index 0 at (0,0,0); `pack_palette` gives
it `I = 0` by its own minimum-squared-error rule, so 23.3's "the bars sit at
RGB (4,4,4)" needed no special case. **0.04 dB** on the Singe window, palette
ceiling unmoved. Black is reserved, not withheld — the mapper may still spend
index 0 on genuinely black pixels; what it buys is that index 0 is black
whatever the scene contains.
~~**THE RE-ENCODE BUNDLE, collected here because it is now four items and they
share one re-measurement.**~~ **DONE, session 28 — FINDINGS 60. Two of the four
closed as NEGATIVES, which is the more useful half.** The four were:
1. ~~reserve palette index 0 as black, `I = 0`~~ **DONE** (23.4, 60.3);
2. ~~`--spans all` as the default~~ **MEASURED AND REFUSED** (E2, 60.4).
267.9 KB/s / 28.88 dB / 2 frames late at `need`, against 448.2 KB/s /
29.07 dB / 1 late at `all`: **+67% of the wire for +0.19 dB and one frame
of 120**, against a container the budget already says is 35% too big.
`need` stays the default. **It was called "the loaded lever on the byte
side" and it is — it is just loaded the wrong way**, and E7 is why. The
GATE container keeps `all`: it is a fixture, not a recipe;
3. ~~re-derive span selection jointly with `lam`~~ **IMPLEMENTED, MEASURED,
NO-OP** (E3, 60.5). `--joint-spans` hands the span pass's freed bytes back
to the lam search and re-spans; all four cells of `{need,all} x
{greedy,joint}` are BYTE-IDENTICAL, and so is `--rc-floor open`. **`lam`
never leaves its floor on any of 120 frames at either floor the encoder
offers** (44.3), so there is nothing to spend the freed bytes on. The code
stays, defaulted off, because a container that moved `lam` off its floor
would make the question live again;
4. ~~**sector-align every record**~~ **DONE — the container is DLX5** (58.3
option C, promoted to a precondition by 59.4; 60.1). Realised cost
**+0.48%** on the wire against the +0.43% predicted, zero clocks, and
**120/120 records start on a sector boundary** where 3/120 did. The disc
and the ring now move the SAME 4,510,208 B and check.sh gates on that
identity, both figures read out of the container instead of written into
the script — the old literals went red on the re-encode, correctly.
**The consumer had to be told too** (60.2): `stream.s` released the ring
to the last byte it READ, which strands up to 511 B of pad a record, and
the ring's own audit caught it on frame 0 while every frame still decoded
pixel-exact. The release rounds to `RECALN` now.
The letterbox no longer gets the palette's closest thing to black — item 1 put
true black at index 0, and `load.i` needed no change, as it said it would not.
~~**P3. A real frame clock.**~~ **DONE, session 22 — FINDINGS 54.**
`src/player/clock.i` derives the tick from the CRTC's own V-DISP through the
MFP, with a remainder-keeping divider whose two constants are read out of the
CRTC at init. **Exactly 12.000000 fps, by construction** — measured at 649 ticks
over 3,000 refreshes where 649.1429 were due, so the remainder still held and
nothing accumulated. **181.35 clocks per V-DISP, 838 per frame, 0.1006% of the
budget**, timed by the 68000 itself because the host's 17.64 ms granularity
cannot see it. `PACEON=0` free-run is untouched and so is the wait loop; the
free-running path executes none of the new code.
The item said "MFP timer or VBL" and **neither can do it alone**: 4e6/12 is not
an integer and no prescale/data pair reaches 12 Hz, while the slowest MFP tick
of any kind is 78.125 Hz; and the raster's 55.4577 Hz has no whole divide near
12 either (4 gives 13.86, 5 gives 11.09). `tools/analysis/23_frame_clock.py`
walks the whole space rather than asserting it.
**What it exposed is bigger than the item.** 12 fps on a 55.4577 Hz raster is
4.6215 refreshes, so a frame gets **4 refreshes (72.13 ms) or 5 (90.16 ms)** and
**there is no 83.33 ms frame** — that figure is the mean slot, and 37.9% of slots
are 13.4% under it. The cadence was ALREADY in every host-paced result in
FINDINGS 49/51, because `stream.lua`'s tick is sampled at frame boundaries and
its gaps were always 4 or 5; nothing had named it. On the gate container it
costs 4 frames of 120 their idle against 1 for the nominal model. **It is not a
dropped frame** — the pace gate lets an overrun eat the next frame's idle and
the clock recovers — but it means every budget in this project is priced against
a slot 37.9% of frames do not get. 54.4.
**Also struck: MAME's raster runs 2.22% fast** (`refresh_mode()` builds the frame
period from `htotal - 8`), so the tree's "1/55.46 s granularity" was 1/56.69 s
throughout. No 68000 cycle figure moves — the CPU clock is unrelated to the
screen — but anything paced by the raster does. 54.5.
**P4. Real transport. P4b DONE, session 26 — FINDINGS 58. P4a DONE at the
transport level, session 27 — FINDINGS 59. What is now between this tree and M2
is THE RE-ENCODE BUNDLE under P2, because the channel refuses a windowed read
(59.4) and 117 of 120 records need one.**
~~Drive the MB89352 instead of a host file.~~ `src/player/scsi.i` selects a SCSI
target and issues READ(10) on the 68000, with no IOCS and no host in the
transfer path: **4,096 B from LBA 0 and 2,048 B from LBA 1000, both byte-exact**
against the host's copy of the same volume.
**This item was listed as blocked and was not.** Session 21 recorded "MAME's
`x68000` has no MB89352 path"; `-exp1 cz6bs1` instantiates one, and FINDINGS
32.4 had read that card's DMA glue back in session 9. The real gap was the 8 KB
`scsiexrom.bin` MAME needs to instantiate the card and **the player never
executes**; a blank placeholder on a separate rompath settles it. **B3 still
wants the real ROM's bytes** and is untouched by this.
**What is left is the half that decides the project**, and it is now two pieces:
~~**P4a. A DMAC configuration that HOLDS THE BUS.**~~ **DONE at the transport
level, session 27 — FINDINGS 59.** `src/player/dma.i` programs HD63450 channel 1
and takes the DATA IN phase: **the same 2,048 B off the disc three ways — PIO,
held, stealing — all three byte-exact.** 57.3's warning was met rather than
worked around: the evidence never reads `$EA0015`. **MTC is sampled by the
instruction after the one that starts the channel, and held it reads zero of
2,048** — the whole transfer happened between two instructions, because the
68000 did not execute in between — against the full count and 426 CPU loop trips
for the stealing configuration. Put the stealing registers in the held slot and
every byte still arrives and the gate goes **red**, which is what says the
counter can come out different (58.3's vacuous-counter trap, avoided
deliberately).
**Three bounds on the apparatus, read out of MAME's source and not inferred**
(59.2): the card has **no request line to the DMAC** (its flow control is
DTACK), so external request — the mode the `W`=5 and `W`=12 rows assume —
cannot be run; **single address** cannot be run either (only channel 0 has
device callbacks); and **only burst is modelled as held**. Of the four rows of
the ladder exactly one, dual address held, has a code path here, and it is the
one demonstrated. The slot pinout has `#EXREQ` at B36, so a real card plausibly
drives it — **that is now B3's sharpest form**.
~~**What is left of P4a is downstream of the container, not of the DMAC**
(59.4)~~**and that block is GONE as of session 28.** `sc_in_data` refused a
windowed read because a channel cannot drop the 300 B in front of a record; the
container is DLX5 now and no record asks for a window. `xfer.i`'s sector
arithmetic already degenerates correctly — `SC_WSKIP` is 0 and `SC_WKEEP` is the
whole record on every one of the 120 — so **what is left of P4a is the wiring:
which loop moves the bytes.** `dma_run.sh`'s windowed-read refusal stays as a
negative control rather than as a description of the container. **This is the
next item, and it is the last one before M2** (60.9).
**P4c (new, and it is a DESIGN CHOICE the tree had not named).** Auto-request is
charged **by time, not by byte** — the channel spends its share of the bus
whether or not a byte is there, so halving the delivery rate DOUBLES the CPU
cost of the same record. The MC68450's GCR sets that share: `BT`/`BR`, four
values, 50/25/12.5/6.25%. `tools/analysis/28_autorequest_cost.py` prices it
against an explicit rate; at 460 KB/s **only the 50% share carries this
container**, at 10.61 clk/B and 47.6% of a frame per record, against 40.4% for
the `W`=9 row and 391.8% measured for PIO. **If B3 comes back saying the real
card drives `#EXREQ`, the ladder applies and this is the fallback; if it does
not, this IS the cost model** and the GCR pair is a number the player has to
choose.
~~**P4b. `scsi.i` behind `ring.i`'s `XF_*` mailbox.**~~ **DONE, session 26 —
FINDINGS 58.** `src/player/xfer.i` answers the mailbox with a real READ(10) per
record: **120 records, 4,488,588 B, pixel-exact, out of the same 256 KB ring,
with a real mid-stream seek in a second pass**. The tiling is the SAME 18 wraps
and 14.7 KB mean hole that 49.4's host producer and 55.4's modelled transport
produced — a third transport, same placement, which is the assertion that
`ring.i` could not tell which side of the seam answered it. The change above the
seam is two `bsr`s, and the one in `ring_seek`'s quiet-wait is not optional:
with the transport inside the machine, that loop is the only thing that can
retire an outstanding request.
**What it cost is the finding, and it re-prices P4a.** `tools/bench/
xfer_cost.sh` subtracts the same 120 frames run twice and gets **87.28 clocks
per delivered byte** — against the 68000's own cycle table for the loop, which
says **87.15**. **0.2% apart**, so it is the instruction stream and not MAME's
device model, and it is therefore the first number this rig has produced that
survives leaving the emulator. At this container's 37,405 B mean record that is
**391.8% of a 12 fps frame**, and the machine's own V-DISP clock agrees from the
other end: **2.57 fps**.
W = 5 single address, bus HELD ............................ 22.4%
W = 9 dual address, held .................................. 40.4%
W = 12 single address, arbitrated .......................... 53.9%
W = 19 dual address, arbitrated -- the IPL ROM's own (52.5) . 85.3%
PIO 87 MEASURED, session 26 ................................ 391.8%
**So P4a is worth 4.6x the worst DMA configuration in this tree and 17.5x the
best**, where before this session it was worth 9 against 19. `W` itself **did
not move by one clock** and is still the largest open number — but what depends
on it just got much larger.
**One more thing P4a inherits (58.3).** A record is not a sector: 117 of 120
start part way into one. PIO absorbs that for free because the CPU is already
touching every byte and simply does not store the ones outside the window — a
property that **disappears the moment the DMAC takes over**, because a channel
writes a contiguous run and cannot drop bytes. The three ways out price as
+1.34% wire and no DMA (windowed PIO), +1.34% wire and **+5 clk/B of copy**
(bounce buffer, which is exactly the cost `aligned` was chosen over `split` to
avoid), or **+0.43% wire and zero clocks** (sector-aligned records in the
container). The last one wins on both axes and is a **re-encode**; see the
bundle under P2. **P4a should be attempted against a sector-aligned container,
not against this one.** *(Session 27: it was, in the only sense that mattered —
the transport now REFUSES the windowed case rather than being trusted not to
reach it, so the bundle is a precondition rather than a plan. 59.4. Session 28:
the container IS one — the realised wire cost is +0.48% against the +0.43%
predicted here, and 120/120 records start on a sector boundary. 60.1.)*
*(original item, still the standing description of the `W` question:)*
Drive the MB89352 instead of a host file. **Session 23
added a second axis to it:** `W` is the clocks stolen per delivered byte, and
55.3 measured that the player's own request loop gives away 3-7% of the pipe
before `W` is even asked about. A transport design has to answer both. The `W`
handshake — clocks stolen per delivered byte, bracketed 5..12 by MC68450 Fig
4-25 — is listed in "Decisions locked" as UNDECIDED and as the thing that
decides the project: `W<=6` fits 0/120 frames, `W=8` misses 47/120. It is a
property of how the player drives the SPC, **so it is ours to choose, not to
receive** (FINDINGS 42.4-42.6). B3 informs it.
**Session 20 promoted this to the project's biggest open number.** FINDINGS 52.5
found the only worked example of a disk DMA configuration on this machine — the
IPL ROM's own — sitting at **16..19 clk/B**, where the whole design fails at any
container size (`15_bus_occupancy.py` sweeps it). The per-byte ladder is 5 clk/B
single-address with the bus held, 9 dual-address held, 12 single-address
arbitrated, 16..19 dual-address arbitrated. **Getting the DMAC to hold the bus
is the difference between 9 and 19**, it is a property of how the player
programs the channel, and demonstrating a configuration that does it is P4's
first job rather than its last.
**Do not quote 42.4's `W <= 6` / `W = 8` sensitivity table for this.** It is in
clocks per WORD and FINDINGS 43 voided it; 52.5 cited it in byte units when
first written and strikes it.
~~**P5. Seek and branch.**~~ **DONE, session 23 — FINDINGS 55.**
`src/player/ring.i` fills the ring on the 68000: `aligned` placement, the
descriptor ring, a prefill policy, 51.2's slack rule as arithmetic the player
can run (`ring_may_seek`), and a seek that quiets the channel and re-addresses
the stream out of the index. It reproduces the host producer's tiling exactly —
18 wraps, 14.7 KB mean hole, pixel-exact — and the host now AUDITS every
placement instead of making it.
The index is a **container change**: DLX4 carries `nframes` u16 record lengths
in the scene header, because `aligned` needs a record's length before it fetches
it and walking the stream is precisely what a player cannot do. Frame payloads
are byte-identical to the DLX3 encode; the scene header goes 5,920 to 6,164 B.
**What it exposed is bigger than the item.** A channel only moves bytes while it
has a request and only the CPU can issue one, so the disc **stands still between
records** by an amount set by the player's loop rather than by the medium — and
no host-filled run could see it. At 488 KB/s in a 256 KB ring, a one-deep
request queue gives away **6.8% of the pipe and underruns 59 of 120 frames**; a
two-deep one gives away 3.4% and underruns none. The container's whole surplus
over the wire at that rate is 8.7%, so the player's own loop was spending most
of the slack 51.3 accumulates. **Prefill is the weaker lever** — six records of
it still leaves 24 underruns at depth 1 — and the fix costs no clocks and no
bytes. 55.3, 55.4.
**P5a (open, and it belongs with P4).** The two-deep queue is modelled as two
mailbox slots. On the machine it is two DMAC channels or one channel with a
chained descriptor array, and which of those is affordable is a `W` question.
**P7. Boot.** The player as an executable loading from the SCSI volume.
Buildable, and empty until P4: there is nothing to boot from yet.
---
## M3 — the vertical slice, and the completion target
**Exit criterion: one decision point, two outcomes, a death clip, with audio,
playing from disc on stock hardware.**
**P6. Audio. P6a DONE, session 34 — FINDINGS 66.** ~~and it is the largest
unpriced risk left in the project.~~ Three quarters of P6 closed in session 33
and the fourth precondition — which decoder the chip runs — closed in 34, on the
machine, through the real DMA channel. ~~**What is left of P6 is P6b (DLXP2, a
container with sound in it) and the refill climb with a second consumer.**~~
**P6 IS CLOSED end of session 38** — P6b in 35, P6c in 36, the level in 37 and
the refill climb with a second consumer in 38 (FINDINGS 70).
*(original framing, kept because every figure below is still the live one:)*
MSM6258 ADPCM, 15.6 kHz mono, **7.8 KB/s**. That figure is in `ratectl.py`'s
budget and nowhere else: not extracted, not encoded, not interleaved into the
container, and **never priced on the bus**. Two reasons to treat it as a risk
rather than a task:
1. A second DMA consumer attacks **the bus** — the resource this project already
established is the binding one, at 86.7% occupied. Clock headroom says
nothing about whether it fits.
2. 7.8 KB/s is a *byte* figure. The last time a byte/word unit error went
unexamined in a delivery budget it cost the project a 2x error in every table
since FINDINGS 5 (session 14, the MB89352 being an 8-bit SPC).
~~**Price it before writing it**: add the ADPCM DMA stream to `15_bus.py` and see
what it does to the 86.7%.~~ **DONE, session 20 — FINDINGS 52.** It is in
`15_bus_occupancy.py` and the answer is **1.25%..1.48% of the frame**, about 4%
of what the decoder leaves. The per-byte cost is no longer a guess borrowed from
the disk: `tools/analysis/21_iplrom_dmac.py` reads the IPL ROM's own HD63450
configuration and finds ch3 dual-address, 8-bit port, cycle steal without hold,
external request — **16..19 clocks per byte**, where `11_cpu_budget.py` had been
charging audio the disk's 5. Both worries above resolve:
1. **The bus concern does not materialise.** A second DMA consumer at 7.8 kB/s
is not what a bus at 88% occupancy is short of.
2. **The unit was checked and is nearly right.** 15.6 kHz = 8 MHz ÷ 512 =
15,625 samples/s, 4 bits each, two to a byte = **7,812.5 B/s exactly**. The
7.8 was decimal kB being multiplied by 1024; 2.4% high, now derived from the
sample rate in `buscost.ADPCM_BYTES_PER_S`.
~~**What is still open in P6 is everything except the bus:** extraction, encode,
container interleave, and what a second stream does to `wire`.~~ **THREE OF THE
FOUR ARE DONE, session 33 — FINDINGS 65.** `tools/encoder/extract_audio.py`
takes the same seconds of the same stream the frames come from;
`tools/encoder/adpcm.py` encodes them (**78,125 B, 21.97 dB**, gated against
ffmpeg's decoder sample-exact by `tools/bench/verify_adpcm.py`);
`tools/analysis/32_audio_wire.py` is the interleave and the wire. **The packed
container's cadence is `F=11, A=14`** — 0.09% padding, 14,336 B held, 582.0 →
**589.6 KB/s** — and the naive one-lump-per-record cadence would have wasted
57.3% of every audio sector. **The codec container pays zero padding**, because
it already has the index the packed one deleted (65.4).
~~**What is left in P6 is the fourth: the refill climb with a second consumer
through a real branch point** (51.3, 55.4). The slack table is in
`32_audio_wire.py` — at 582.0 KB/s exactly, silent breaks even and sounded
starves — but a branch point has not been run with audio on the wire.~~
**DONE, session 38 — FINDINGS 70**, and it produced three things the slack table
could not: the climb multiplier (up to **3.30x** for **1.7%** of the wire), the
fact that the packed branch has **no accumulator to climb**, and the group-entry
silence (**mean 416.5 ms** at 409 of the game's 612 branch points) that reopens
the cadence pick on a third column. `tools/analysis/36_branch_audio.py`.
**P6a. WHICH DELTA FORMULA DOES THE MSM6258 RUN? (new, session 33, FINDINGS
65.2 — and it is a precondition, not a refinement.)** ffmpeg computes
`((2*(n&7)+1) * step) >> 3`; the OKI datasheet's form truncates per term. They
differ by **at most 3 in 12-bit units per sample** and, because ADPCM is
recursive, **encoding for one and decoding on the other costs 25 dB — the noise
comes out louder than the signal.** No encoder output can be committed to a
container before this is answered.
**It does not need a board.** MAME's x68000 HAS the chip: `:okim6258`, with the
68000 reaching it at **`$E92001` and `$E92003`** — both read out of the machine's
own program map by `tools/bench/probe_adpcm.lua`, not from folklore. What did
NOT work is feeding it from Lua: `probe_adpcm2.lua` / `probe_adpcm3.lua` swept
control 0..3 against PPI port C 0..15 and `-wavwrite` recorded silence
throughout (65.5). The gap is the register semantics, and the way to close it is
**from 68000 code with the IPL ROM's own channel-3 DMAC configuration**, which
`21_iplrom_dmac.py` already reads out of the ROM — the one ADPCM path in this
machine that is known-correct because Sharp wrote it. That is also the real
design, so it is not scaffolding.
**P6b. A CONTAINER WITH SOUND IN IT. DONE, session 35 — FINDINGS 67.**
~~DLXP1 has no audio section. 65.3 is the arithmetic a DLXP2 is built from and no
byte of one is written.~~ DLXP2 is written, gated and carried by the player:
`tools/encoder/dlxp.py`, `pack.py --audio`, `tools/analysis/34_packed_audio.py`,
the third LBA term in `src/player/packed.s`. The cadence arithmetic came through
unchanged (F=11, A=14, 589.6 KB/s) and the **payload** did not: a lump's audio is
7,161 or 7,162 B of a 7,168 B sector run, and feeding the chip the whole lump
drifts 1.25 s over the game.
**P6c. AUDIO OUT OF THE CONTAINER, ON THE MACHINE. DONE, session 36 — FINDINGS
68.** The lump buffer is `PG_ABUF` (three slots, prefill two, and the minimum
depth is still unmeasured), the remainder accumulator is `pg_apay`, and the
service runs from inside `dma.i`'s transfer wait through `DM_HOOK` because once
a frame is a 90 ms seam by construction. **Two channels have run at once**: the
bytes are identical in both DMAC configurations and the sound is not — held
costs 236 ms of replayed byte against stealing's 0.51 ms. ~~Every
piece exists and none is joined up. The container carries the bytes (67); the
transport is `src/player/adpcm.i`'s channel-3 configuration, which is the IPL
ROM's own and worked first time (66.1). What is missing is **the lump buffer**
— 14,336 B, double-buffered, allocated by nobody — and **the remainder
accumulator**, which is three instructions and is not optional.~~
**And the interaction neither half has met: TWO CHANNELS AT ONCE.** The video
channel holds the bus and halts the 68000, which already costs the frame clock
half its ticks without the clock being able to tell (64.3). An audio channel
that must be serviced *during* that hold has never been run. 52's 1.25%..1.48%
of a frame is a figure measured in isolation, and `32_audio_wire.py` names what
it turns into here: audio does not merely cost clocks, it costs **darkness**.
**E6. Container v2** — audio interleave, per-record index, scene table. Depends
on P6's answer and on P5's index.
**E7. A BYTE TARGET, AND IT COMES OUT OF THE BUS BUDGET RATHER THAN OUT OF
TASTE (new, session 27, FINDINGS 59.7; re-measured session 28, 60.7).** The
frame affords **6.69 clocks a byte** after the measured decode and the audio,
and a dual-address byte costs **9**. So *if* B3 comes back saying the card
cannot drive `#EXREQ`, the container has to reach **27,924 B a frame — 327 KB/s
of payload** to fit at 12 fps, where the DLX5 gate container delivers 37,585 B
and 440 KB/s: **35% too big.** (It was 6.74 / 328 / 34% against the DLX4
container. **The bundle moved the target by one KB/s and moved no conclusion**,
which is what a precondition is supposed to do.)
Three things make this less alarming than the number looks, and one makes it
worse:
- The gate container is **deliberately the heaviest thing the encoder emits**
(span-heavy, the 488 recipe, every block mode exercised). It is a test
fixture, not a shipping target.
- A lighter container **also decodes cheaper**, so the 68.5% decode term falls
with the byte term. 328 KB/s is the pessimistic reading of the lever.
- `rc_fr_singe_scsi_cpufit.dlx` already exists — the encoder has had a
CPU-fitting mode since session 11.
- **Worse:** `15_bus_occupancy.py` REFUSES to price the cpufit container,
correctly, because the C68K measurement it cross-checks against belongs to
the gate container. **So E7 starts with a harness re-run**
(`tools/bench/c68k/run.sh`) against whichever container is to be the target,
and until that is done "34% too big" is a statement about the fixture rather
than about the project.
~~**G1. Import the scene graph — early, because it is a measurement input.**~~
**DONE, session 24 — FINDINGS 56.** It was pulled ahead for exactly the reason
given, and it paid: **the worst gap between two consecutive decision points is
zero**, and 5.4% of the game's 612 branch transitions are. Two seeks can fall
back to back with no play between them, so 51.2's slack rule can be answered NO
by the content rather than by the buffer.
It does not break the design — a branch on an empty ring costs the prefill
(149.7 ms, 1.80 frame slots at 488 KB/s), not the climb — but it removes the
margin: at 488 KB/s in a 256 KB ring, **76% of this game's branch points arrive
before the ring has refilled**, and a 512 KB ring makes that 90%. **The ring is
not the lever; the surplus is.**
Two constraints on the input layer came with it: the arcade needs **eight
directions**, and the shortest input window is **98 ms** against a 72.13/90.16 ms
frame slot, so input cannot be polled on the frame tick (56.7).
**The cross-check plan was wrong and is struck.** The SNES chapters are
*derived* from DirkSimple, by their own README, so there is one transcription and
not two; the diff catches conversion errors only (56.2). **Nothing is vendored:**
`tools/import/scenegraph.py` is the one file coupled to those projects and it
writes this project's own `DLXSCENE1` schema into gitignored `tmp/`
(USER DECISION, session 24).
---
## M4 — the whole game
Listed for completeness; past M3 these are scope, not risk.
- **C1. Full-disc survey**, 22.8 minutes. Classify **content / menu / bonus**
not menu vs content: the two largest streams are bonus material and look like
content by size, duration and bitrate alike (25.1). Run
`07_motion_survey.py` per stream first for a hot-window shortlist.
**Gated by E4, and parked with it (session 29).**
- ~~**E4. `H.build` k-means**, 51 s of a 55 s run, once per scene.~~ **PARKED,
session 29 (USER DECISION).** It was the thing to attack before C1. It builds
**VQ codebooks**, and a decoder-free packed player has no VQ — so this is
encoder work on the branch that is no longer being built on. It comes back if
and only if B2 goes MAME's way. C1 is gated by it and is parked with it.
- ~~**E2. `--spans all` as default.**~~ **MEASURED AND REFUSED, session 28 —
60.4.** It was "a recommendation, not a measurement" since 43.6.1 and it is a
measurement now: **+67% of the wire for +0.19 dB and one frame of 120.** It
IS the loaded lever on the byte side (44.3) — it is loaded the wrong way, and
E7 is why. What the entry predicted is exactly what it does: it raises `wire`,
which shrinks `pipe - wire`, which lengthens the refill climb. `need` stays
the default; the GATE container keeps `all` because it is a fixture.
- ~~**E3. Re-derive span selection jointly with `lam`** (39.3).~~
**IMPLEMENTED, MEASURED, NO-OP, session 28 — 60.5.** `--joint-spans` emits
byte-identical containers in all four `{need,all} x {greedy,joint}` cells and
at both lam floors, because **`lam` never leaves its floor on any of 120
frames**. Kept and defaulted off.
- **C2. Framing** — crop vs squash vs wide (FINDINGS 12). Needs an eyeball
against arcade reference, not a measurement. Cheap; blocks only final encodes.
- **C3. Disk image packaging**, ~1.09 GiB at the candidate rate.
- **G2/G3.** Branching, input windows, death clips, attract mode; playtest.
---
## Dependency summary
```
P4a DONE (59): the channel drives the data phase and
holds the bus -- 391.7% of a frame becomes 40..95%
B3 #EXREQ? ──┬─ YES ─> single address, 5 clk/B, 92.4% ── FITS ──┐
│ │
└─ NO ──> auto-request, 9 clk/B FLOOR, 110.4% ──> E7 byte target
(the frame affords 6.69; 59.7, 60.7) 327 KB/s
P2 re-encode bundle DONE (60): DLX5, records ARE sectors ─────────┤
E2 refused on measurement, E3 a no-op -- 60.4, 60.5 │
E7/E4/C1 PARKED session 29: encoder work waits on B2 ───────────┤
K1 palette-register DMA? ─> K2 packed container ─> K3 end to end ──┤
K1 DONE s30 (62), K2 DONE s31 (63), K3 DONE s32 (64): │
120/120 pixel-exact, both palette orders, off a real volume │
└─> and K3 found that the WRITE WINDOW IS THE FRAME, so: │
K4 (DMA to RAM + the measured 27.3% paint) is the player │
that is ON SCREEN below a 2,131 KB/s BURST rate -- which │
is 3.7x the wire, so below every rate anyone has proposed │
P4a WIRING (the channel behind ring.i's mailbox) <- THE LAST ITEM ─┤
P1 P2 P3 P4b P5 P7, P6 bus cost (52), G1 scene graph (56) ────────┼─> M2 ─>
B1 seek+rate (sets HEADROOM, not fit) ─────────────────────────────┘ M3 ─> M4
B2 blanking ─┬─ NOT blanked ─> K3's DMAC-DIRECT player is the one: 54.9% of a
│ frame at the 9 clk/B floor against the codec's
│ 110.4%, on screen the whole slot, and K4's paint
│ would be 27.3% spent on nothing (61.4, 61.5)
└─ blanked ──────> K3's player is on screen for
1 - record/(BURST x slot) of every slot, which is
ZERO at the container's own wire -- so K4 is the
player, at 82.2% of a frame and 99,328 B of RAM
(64.2). Neither answer kills the branch and each
picks a different player.
B1 BURST rate (NEW, 64.2) ──> which of the two K3/K4 wins, if B2 blanks
P6 audio: encoder gated (65), P6a the chip's own decoder measured (66),
P6b DLXP2 written and gated, 589.6 KB/s on the wire (67),
P6c PLAYED -- 78,125 B to the chip beside the video channel (68)
└─> and the second consumer is now ANOTHER INPUT TO B1/B2:
stealing -> 0.51 ms of seam over 10 s \ same bytes,
held -> 236 ms, every lump boundary / different sound
└─> what is LEFT of P6: the LEVEL (66.3, oldest open item) and the
refill climb with the second consumer through a branch point
```
**Read that top-left branch as the project's live question.** Everything else
on the diagram is work; `#EXREQ` is a fact about a board nobody here has, and it
decides which of the two lower paths the player is on.
## Standing rules that apply to all of it
- **Green light first and last.** `./tools/bench/check.sh`, ALL GREEN, before and
after. **Never two MAME jobs at once** — session 18 did it, two `decode.lua`
runs shared a log file, and it produced a 0-byte log and 15 wasted minutes.
- **Name the layer.** Emulated, or real hardware. Every progress claim.
- **Label measured / estimated / folklore.** A rate with no provenance is
folklore even when it is plausible, and this project has already paid for that
twice.
- **No new default constants.** Rates stay explicit arguments. If a measurement
is not available, report the sensitivity across several rates rather than
picking one.