FINDINGS 59.7. tools/analysis/15_bus_occupancy.py has always answered "what does each W cost" and never "what can the frame afford", and after 59.2 those are not the same question. It now answers both, and takes an optional --kbps for the auto-request rows -- the only rows whose cost depends on how long the record takes to arrive. On the gate container at 12 fps, decode term MEASURED: decode 68.5%, audio 1.25%, HEADROOM 30.2% = 6.74 clocks per byte at a 37,403 B record. Against that, P4a cut the transport from 391.7% of a frame to 40..95% -- four to ten times, the largest movement in this project's cost model since the decoder was written -- and it still does not fit. A dual-address byte is a 4-clock read of the device plus a 5-clock write to memory, so 9 clk/B is a FLOOR and the frame affords 6.74. No GCR share goes under it and no delivery rate goes under it: a share decides whether the channel sits at the floor or above it. At 460 KB/s max-rate totals 165.1% and LRAR at 50% totals 117.4%, and a 50% share tops out at 543 KB/s, above which the channel is the bottleneck and the rate falls back to exactly that floor. So 59.2's three bounds arrive in the budget as one sentence: the configurations this machine can run are the ones the frame cannot afford, and the one it can afford -- single address, 5 clk/B, 92.2% total, 7.8% spare -- needs the device to ACK the DMAC, which needs a request line MAME does not connect and the slot pinout does have at B36/B37. ROADMAP re-ranks accordingly. B3 stops being a constant to look up and becomes DOES THE CARD DRIVE #EXREQ, ahead of B1: B1 sets how much headroom the player has, B3 decides whether there is any. New E7 carries the other branch -- if the answer is no, the container must reach 27,995 B a frame, 328 KB/s of payload, against 438 now. The dependency diagram is redrawn around that fork. The scope is stated rather than buried: this is the GATE container, deliberately the heaviest thing the encoder emits, and the lighter cpufit family was NOT priced -- 15_bus_occupancy.py refuses it, correctly, because the C68K measurement it cross-checks against belongs to the gate container. E7 therefore begins with a harness re-run, and until then "34% too big" is a statement about the fixture and not about the project. check.sh ALL GREEN before and after. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
30 KiB
Roadmap — remaining work to a completion target
Written end of session 19 (2026-08-24), against a tree that is ALL GREEN. Amended end of session 21: P1 done, P2 half done (FINDINGS 53). Amended end of session 22: P3 done (FINDINGS 54). Amended end of session 23: P5 done (FINDINGS 55). Amended end of session 24: G1 done (FINDINGS 56). Amended end of session 25: P4 HALF done (FINDINGS 57). Amended end of session 26: P4b done, P4a is the last open item before M2 (FINDINGS 58). Amended end of session 27: P4a done at the transport level; THE RE-ENCODE BUNDLE under P2 is now the only thing between this tree and M2, because 59.4 made sector-aligned records a precondition the transport enforces rather than a preference (FINDINGS 59). And 59.7 re-ranks what is left: the frame affords 6.74 clocks a byte, a dual-address byte costs 9, so B3 stopped being a constant to look up and became the question of whether the design fits at all.
THE COMPLETION TARGET IS M3, THE VERTICAL SLICE (USER DECISION): one scene tree — a decision point, two outcomes, a death clip — with audio, streaming from a real SCSI volume on a stock 2 MB machine, playable. That is the point at which every layer of this design has been shown to work at once. M4 is listed because it is real work, but past M3 it is content grinding rather than open questions.
docs/STATUS.md remains the session-by-session record and the handoff. This file
is the shape of what is left; where the two disagree about what is done, STATUS
is the one with the measurements and this one is the one that goes stale. Both
were wrong about two encoder gaps until this file was written — see "What was
already done" below.
Status of the four resources
The project's own framing, restated because every item below is priced in one of these units:
| resource | state |
|---|---|
| 68000 local bus | the binding one. Decoder occupies 86.7%; 52 of 53 missed frames miss on the bus, not the clock (FINDINGS 38). |
| 68000 clocks | measured, and the rate controller binds on them. |
| Delivery rate | no working figure, deliberately (FINDINGS 50, USER DECISION). Every tool REQUIRES an explicit rate. |
| Seek time | no figure at all, and never had one. 51.3/51.4 made it matter. |
| W, clocks stolen per delivered byte | 5 single-address held, 9 dual held, 12 single arbitrated; the IPL ROM's own disk channel is 16..19 (52.5). The largest open number in the project. Session 27 added the row underneath it: with no external request line on the card (59.2) the channel is auto-requested and is charged by time rather than by byte, so at 460 KB/s a 50% bus share costs 10.61 clk/B and a smaller share cannot carry the rate at all (59.3). |
| The frame's headroom for a transport | 6.74 clk/B — 30.2% of a 12 fps frame, after the MEASURED decode (68.5%) and best-case audio (1.25%), at the gate container's 37,403 B record. New in session 27, 59.7, and it is the number every row above is now read against. A dual-address byte's floor is 9. |
What was already done, and was still on the list
Found while inventorying for this file. Both had been closed in code for several
sessions and were still listed as open gaps in docs/STATUS.md:
- 4-byte record padding.
DLX2,encode.py:139-156, inside rate-control accounting, reported per frame and per second. - CPU cost in the mode decision.
vq_hybrid.py:218, priced against measured per-mode cycles with the exact clustered SKIP rule.
Both entries are now struck in STATUS. The lesson is procedural: a gap list that is only ever appended to manufactures phantom work. Anything crossed off below should be crossed off in STATUS in the same sitting.
Blocked on hardware this tree does not have
None of these block M2 or M3 software work, because session 18 forced every rate to be an explicit argument. They set constants, and two of them decide how much headroom the finished player has.
B1. Measure the BlueSCSI — throughput AND seek time.
Throughput has an acceptance test already derived from real record sizes:
513.2 KB/s for the session-14 candidate, 451.4 KB/s for the gate
container (19_ring_stream.py, FINDINGS 49.5). Seek time has nothing.
51.3/51.4 is why the second half matters: slack is accumulated out of
pipe - wire, so what a branch point costs is set by the rate and the time since
the last branch, not by the ring size. At 460 KB/s every ring from 192 KB to
512 KB is rate-bound and never fills. Do not substitute a guess — run at
several explicit rates and report the sensitivity. That is exactly how the
retired pipe constant survived five sessions after 42.1 called it folklore.
B2. Does buffer mode blank the display? probe_bit11_blank.lua is written
and settles it in minutes on a real board. FINDINGS 48 shifted the prior toward
MAME and toward "unusable" — do not pre-build on 1.0 B/pixel. Same sitting:
the priority register 0xE82500 at 0x0000 (47.3).
B3. Single-address vs dual-address DMA — and now, first, DOES THE CARD DRIVE
#EXREQ? 242 KB/s and 0.69 dB was the old framing. Session 27 promoted this
to the item that decides whether the design fits at all (59.7). The frame
affords 6.74 clocks a byte; a dual-address byte is 9 — a 4-clock read of
the device plus a 5-clock write to memory — so no dual-address configuration
fits this container at 12 fps, at any delivery rate and any GCR share. Single
address is 5 clk/B and fits at 92.2% of the frame with 7.8% to spare, and it
needs the device to ACK the DMAC directly, which needs the request line.
MAME cannot answer it: x68k_scsiext.cpp's drq_w only stores a flag and
the expansion slot has no request path to the HD63450 at all (59.2). The slot
PINOUT has #EXREQ at B36 and #EXACK at B37, so the provision exists on the
real board. What is wanted from hardware is therefore narrower and sharper
than "disassemble the ROM": does a CZ-6BS1 assert #EXREQ during a data phase,
and will the HD63450 run a single-address channel against it? A scope or a
logic analyser answers the first; the second is a program the player can run.
Sourcing scsiexrom.bin (8 KB, CRC 7be488de) and disassembling its DMAC init
is still the cheapest paper route to the same answer, because whatever Sharp's
own driver programs into DCR's DTYP is a statement about what the card supports.
FINDINGS 48.4. Not on this machine (checked, session 18).
Ranking, stated because B1 used to outrank this: B1 (throughput and seek)
sets how much headroom the finished player has. B3 decides whether there is
any. If the card drives #EXREQ, the ladder applies and the design fits with
room. If it does not, the fallback is limited-rate auto-request at a share the
player picks (P4c), and the container has to come down to 328 KB/s of payload
to fit at the 9 clk/B floor — 34% below where the gate container sits (59.7).
Session 20 moved the prior hard, and it moved the wrong way (FINDINGS 52.5). The IPL ROM is on this machine, and
tools/analysis/21_iplrom_dmac.pyreads its HD63450 setup: the on-board disk channel (ch1, SASI) isDCR = $80— dual address, 8-bit port, cycle steal WITHOUT hold, withREQG = 10external request, i.e. a full bus arbitration per byte. That is 16..19 clocks per delivered byte, above the whole 5..12 bracket 42.4 costs P4 in. Same vendor, same DMAC, same class of 8-bit port — but it is notscsiexrom.bin, so B3 stays open. What it changes is that a cheap configuration is now the thing that has to be shown, not assumed.
M2 — a player, as opposed to a decoder
decode.s draws pixel-exact frames from RAM Lua pre-loaded; stream.s decodes
out of a bounded ring fed by a host file on a paced clock. Neither is a player.
Exit criterion: boots from a real SCSI volume on a stock 2 MB X68000, plays one scene at 12 fps from disc, no host-file pipe, no Lua in the loop. Silent.
P1. Codebook expansion on the 68000. DONE, session 21 — FINDINGS 53.
src/player/load.i expands both codebooks out of the raw container header,
byte-exact against dlxload.py on both CPU cores. 9.26 ms, and it was
priced where it lands rather than treated as free setup: the scene header is
5,920 B that no rate table in this tree counted, and in the currency of
51.3 — accumulated slack — those bytes lengthen the refill climb by 138 ms at
488 KB/s and by 1.099 s at 451.4 KB/s, because the surplus they are divided
by goes to zero. The whole fixed cost of a scene change is about a third of one
frame slot; what makes a branch point expensive is still the seek and the climb.
Shipping the codebooks pre-expanded was considered and refused: it trades
9.26 ms of CPU for 5,120 more header bytes, which is a wash in milliseconds and
not a wash in kind (53.6).
P2. Palette packing on the 68000. HALF DONE, session 21 — FINDINGS 53.
The encoder still emits RGB888; the X68000 word packing is Lua-side. The
packing is on the 68000: pal_pack writes 256 words straight into $E82000
with I chosen per entry by minimum squared error (1.96 dB, 23.3), gated on
the words read back out of the palette registers. 9.70 ms per scene, plus
5.29 ms of scene-independent table build hoisted to boot (53.3).
What is left is the other half of the sentence: reserve index 0 as black with
I = 0 (23.4). That is ENCODER-side, it changes the container, and it moves
every constant fitted to the gate container, so it is a re-encode plus a
re-measurement rather than an edit.
THE RE-ENCODE BUNDLE, collected here because it is now four items and they share one re-measurement. Nothing below is worth a container revision on its own; together they are one:
- reserve palette index 0 as black,
I = 0(23.4, this item); --spans allas the default (E2, and it is the loaded lever on the byte side);- re-derive span selection jointly with
lam(E3); - sector-align every record (58.3; PROMOTED TO A PRECONDITION in session
27, 59.4 —
sc_in_datarefuses a windowed read under the DMAC, so this is what the channel is waiting for) — +0.43% on the wire, zero clocks, and it is what lets P4a's DMA channel write straight into the ring with no window and no bounce copy. Until then the letterbox gets the palette's closest thing to black (index 255 on the gate container);load.ireports whichever index that is and needs no change when it becomes 0.
P3. A real frame clock. DONE, session 22 — FINDINGS 54.
src/player/clock.i derives the tick from the CRTC's own V-DISP through the
MFP, with a remainder-keeping divider whose two constants are read out of the
CRTC at init. Exactly 12.000000 fps, by construction — measured at 649 ticks
over 3,000 refreshes where 649.1429 were due, so the remainder still held and
nothing accumulated. 181.35 clocks per V-DISP, 838 per frame, 0.1006% of the
budget, timed by the 68000 itself because the host's 17.64 ms granularity
cannot see it. PACEON=0 free-run is untouched and so is the wait loop; the
free-running path executes none of the new code.
The item said "MFP timer or VBL" and neither can do it alone: 4e6/12 is not
an integer and no prescale/data pair reaches 12 Hz, while the slowest MFP tick
of any kind is 78.125 Hz; and the raster's 55.4577 Hz has no whole divide near
12 either (4 gives 13.86, 5 gives 11.09). tools/analysis/23_frame_clock.py
walks the whole space rather than asserting it.
What it exposed is bigger than the item. 12 fps on a 55.4577 Hz raster is
4.6215 refreshes, so a frame gets 4 refreshes (72.13 ms) or 5 (90.16 ms) and
there is no 83.33 ms frame — that figure is the mean slot, and 37.9% of slots
are 13.4% under it. The cadence was ALREADY in every host-paced result in
FINDINGS 49/51, because stream.lua's tick is sampled at frame boundaries and
its gaps were always 4 or 5; nothing had named it. On the gate container it
costs 4 frames of 120 their idle against 1 for the nominal model. It is not a
dropped frame — the pace gate lets an overrun eat the next frame's idle and
the clock recovers — but it means every budget in this project is priced against
a slot 37.9% of frames do not get. 54.4.
Also struck: MAME's raster runs 2.22% fast (refresh_mode() builds the frame
period from htotal - 8), so the tree's "1/55.46 s granularity" was 1/56.69 s
throughout. No 68000 cycle figure moves — the CPU clock is unrelated to the
screen — but anything paced by the raster does. 54.5.
P4. Real transport. P4b DONE, session 26 — FINDINGS 58. P4a DONE at the
transport level, session 27 — FINDINGS 59. What is now between this tree and M2
is THE RE-ENCODE BUNDLE under P2, because the channel refuses a windowed read
(59.4) and 117 of 120 records need one.
Drive the MB89352 instead of a host file. src/player/scsi.i selects a SCSI
target and issues READ(10) on the 68000, with no IOCS and no host in the
transfer path: 4,096 B from LBA 0 and 2,048 B from LBA 1000, both byte-exact
against the host's copy of the same volume.
This item was listed as blocked and was not. Session 21 recorded "MAME's
x68000 has no MB89352 path"; -exp1 cz6bs1 instantiates one, and FINDINGS
32.4 had read that card's DMA glue back in session 9. The real gap was the 8 KB
scsiexrom.bin MAME needs to instantiate the card and the player never
executes; a blank placeholder on a separate rompath settles it. B3 still
wants the real ROM's bytes and is untouched by this.
What is left is the half that decides the project, and it is now two pieces:
P4a. A DMAC configuration that HOLDS THE BUS. DONE at the transport
level, session 27 — FINDINGS 59. src/player/dma.i programs HD63450 channel 1
and takes the DATA IN phase: the same 2,048 B off the disc three ways — PIO,
held, stealing — all three byte-exact. 57.3's warning was met rather than
worked around: the evidence never reads $EA0015. MTC is sampled by the
instruction after the one that starts the channel, and held it reads zero of
2,048 — the whole transfer happened between two instructions, because the
68000 did not execute in between — against the full count and 426 CPU loop trips
for the stealing configuration. Put the stealing registers in the held slot and
every byte still arrives and the gate goes red, which is what says the
counter can come out different (58.3's vacuous-counter trap, avoided
deliberately).
Three bounds on the apparatus, read out of MAME's source and not inferred
(59.2): the card has no request line to the DMAC (its flow control is
DTACK), so external request — the mode the W=5 and W=12 rows assume —
cannot be run; single address cannot be run either (only channel 0 has
device callbacks); and only burst is modelled as held. Of the four rows of
the ladder exactly one, dual address held, has a code path here, and it is the
one demonstrated. The slot pinout has #EXREQ at B36, so a real card plausibly
drives it — that is now B3's sharpest form.
What is left of P4a is downstream of the container, not of the DMAC (59.4):
sc_in_data refuses a windowed read when the data phase is the channel's,
because a channel writes a contiguous run and cannot drop the 300 B in front of
a record. So putting the channel behind ring.i's mailbox waits on the
re-encode bundle.
P4c (new, and it is a DESIGN CHOICE the tree had not named). Auto-request is
charged by time, not by byte — the channel spends its share of the bus
whether or not a byte is there, so halving the delivery rate DOUBLES the CPU
cost of the same record. The MC68450's GCR sets that share: BT/BR, four
values, 50/25/12.5/6.25%. tools/analysis/28_autorequest_cost.py prices it
against an explicit rate; at 460 KB/s only the 50% share carries this
container, at 10.61 clk/B and 47.6% of a frame per record, against 40.4% for
the W=9 row and 391.8% measured for PIO. If B3 comes back saying the real
card drives #EXREQ, the ladder applies and this is the fallback; if it does
not, this IS the cost model and the GCR pair is a number the player has to
choose.
P4b. DONE, session 26 —
FINDINGS 58. scsi.i behind ring.i's XF_* mailbox.src/player/xfer.i answers the mailbox with a real READ(10) per
record: 120 records, 4,488,588 B, pixel-exact, out of the same 256 KB ring,
with a real mid-stream seek in a second pass. The tiling is the SAME 18 wraps
and 14.7 KB mean hole that 49.4's host producer and 55.4's modelled transport
produced — a third transport, same placement, which is the assertion that
ring.i could not tell which side of the seam answered it. The change above the
seam is two bsrs, and the one in ring_seek's quiet-wait is not optional:
with the transport inside the machine, that loop is the only thing that can
retire an outstanding request.
What it cost is the finding, and it re-prices P4a. tools/bench/ xfer_cost.sh subtracts the same 120 frames run twice and gets 87.28 clocks
per delivered byte — against the 68000's own cycle table for the loop, which
says 87.15. 0.2% apart, so it is the instruction stream and not MAME's
device model, and it is therefore the first number this rig has produced that
survives leaving the emulator. At this container's 37,405 B mean record that is
391.8% of a 12 fps frame, and the machine's own V-DISP clock agrees from the
other end: 2.57 fps.
W = 5 single address, bus HELD ............................ 22.4%
W = 9 dual address, held .................................. 40.4%
W = 12 single address, arbitrated .......................... 53.9%
W = 19 dual address, arbitrated -- the IPL ROM's own (52.5) . 85.3%
PIO 87 MEASURED, session 26 ................................ 391.8%
So P4a is worth 4.6x the worst DMA configuration in this tree and 17.5x the
best, where before this session it was worth 9 against 19. W itself did
not move by one clock and is still the largest open number — but what depends
on it just got much larger.
One more thing P4a inherits (58.3). A record is not a sector: 117 of 120
start part way into one. PIO absorbs that for free because the CPU is already
touching every byte and simply does not store the ones outside the window — a
property that disappears the moment the DMAC takes over, because a channel
writes a contiguous run and cannot drop bytes. The three ways out price as
+1.34% wire and no DMA (windowed PIO), +1.34% wire and +5 clk/B of copy
(bounce buffer, which is exactly the cost aligned was chosen over split to
avoid), or +0.43% wire and zero clocks (sector-aligned records in the
container). The last one wins on both axes and is a re-encode; see the
bundle under P2. P4a should be attempted against a sector-aligned container,
not against this one. (Session 27: it was, in the only sense that mattered —
the transport now REFUSES the windowed case rather than being trusted not to
reach it, so the bundle is a precondition rather than a plan. 59.4.)
(original item, still the standing description of the W question:)
Drive the MB89352 instead of a host file. Session 23
added a second axis to it: W is the clocks stolen per delivered byte, and
55.3 measured that the player's own request loop gives away 3-7% of the pipe
before W is even asked about. A transport design has to answer both. The W
handshake — clocks stolen per delivered byte, bracketed 5..12 by MC68450 Fig
4-25 — is listed in "Decisions locked" as UNDECIDED and as the thing that
decides the project: W<=6 fits 0/120 frames, W=8 misses 47/120. It is a
property of how the player drives the SPC, so it is ours to choose, not to
receive (FINDINGS 42.4-42.6). B3 informs it.
Session 20 promoted this to the project's biggest open number. FINDINGS 52.5
found the only worked example of a disk DMA configuration on this machine — the
IPL ROM's own — sitting at 16..19 clk/B, where the whole design fails at any
container size (15_bus_occupancy.py sweeps it). The per-byte ladder is 5 clk/B
single-address with the bus held, 9 dual-address held, 12 single-address
arbitrated, 16..19 dual-address arbitrated. Getting the DMAC to hold the bus
is the difference between 9 and 19, it is a property of how the player
programs the channel, and demonstrating a configuration that does it is P4's
first job rather than its last.
Do not quote 42.4's W <= 6 / W = 8 sensitivity table for this. It is in
clocks per WORD and FINDINGS 43 voided it; 52.5 cited it in byte units when
first written and strikes it.
P5. Seek and branch. DONE, session 23 — FINDINGS 55.
src/player/ring.i fills the ring on the 68000: aligned placement, the
descriptor ring, a prefill policy, 51.2's slack rule as arithmetic the player
can run (ring_may_seek), and a seek that quiets the channel and re-addresses
the stream out of the index. It reproduces the host producer's tiling exactly —
18 wraps, 14.7 KB mean hole, pixel-exact — and the host now AUDITS every
placement instead of making it.
The index is a container change: DLX4 carries nframes u16 record lengths
in the scene header, because aligned needs a record's length before it fetches
it and walking the stream is precisely what a player cannot do. Frame payloads
are byte-identical to the DLX3 encode; the scene header goes 5,920 to 6,164 B.
What it exposed is bigger than the item. A channel only moves bytes while it has a request and only the CPU can issue one, so the disc stands still between records by an amount set by the player's loop rather than by the medium — and no host-filled run could see it. At 488 KB/s in a 256 KB ring, a one-deep request queue gives away 6.8% of the pipe and underruns 59 of 120 frames; a two-deep one gives away 3.4% and underruns none. The container's whole surplus over the wire at that rate is 8.7%, so the player's own loop was spending most of the slack 51.3 accumulates. Prefill is the weaker lever — six records of it still leaves 24 underruns at depth 1 — and the fix costs no clocks and no bytes. 55.3, 55.4.
P5a (open, and it belongs with P4). The two-deep queue is modelled as two
mailbox slots. On the machine it is two DMAC channels or one channel with a
chained descriptor array, and which of those is affordable is a W question.
P7. Boot. The player as an executable loading from the SCSI volume. Buildable, and empty until P4: there is nothing to boot from yet.
M3 — the vertical slice, and the completion target
Exit criterion: one decision point, two outcomes, a death clip, with audio, playing from disc on stock hardware.
P6. Audio — and it is the largest unpriced risk left in the project.
MSM6258 ADPCM, 15.6 kHz mono, 7.8 KB/s. That figure is in ratectl.py's
budget and nowhere else: not extracted, not encoded, not interleaved into the
container, and never priced on the bus. Two reasons to treat it as a risk
rather than a task:
- A second DMA consumer attacks the bus — the resource this project already established is the binding one, at 86.7% occupied. Clock headroom says nothing about whether it fits.
- 7.8 KB/s is a byte figure. The last time a byte/word unit error went unexamined in a delivery budget it cost the project a 2x error in every table since FINDINGS 5 (session 14, the MB89352 being an 8-bit SPC).
Price it before writing it: add the ADPCM DMA stream to DONE, session 20 — FINDINGS 52. It is in
15_bus.py and see
what it does to the 86.7%.15_bus_occupancy.py and the answer is 1.25%..1.48% of the frame, about 4%
of what the decoder leaves. The per-byte cost is no longer a guess borrowed from
the disk: tools/analysis/21_iplrom_dmac.py reads the IPL ROM's own HD63450
configuration and finds ch3 dual-address, 8-bit port, cycle steal without hold,
external request — 16..19 clocks per byte, where 11_cpu_budget.py had been
charging audio the disk's 5. Both worries above resolve:
- The bus concern does not materialise. A second DMA consumer at 7.8 kB/s is not what a bus at 88% occupancy is short of.
- The unit was checked and is nearly right. 15.6 kHz = 8 MHz ÷ 512 =
15,625 samples/s, 4 bits each, two to a byte = 7,812.5 B/s exactly. The
7.8 was decimal kB being multiplied by 1024; 2.4% high, now derived from the
sample rate in
buscost.ADPCM_BYTES_PER_S.
What is still open in P6 is everything except the bus: extraction, encode,
container interleave, and what a second stream does to wire — and therefore to
pipe - wire, and therefore to 51.3's refill climb. That last one is the
interaction to price next, and it is E2's question with a second consumer in it.
E6. Container v2 — audio interleave, per-record index, scene table. Depends on P6's answer and on P5's index.
E7. A BYTE TARGET, AND IT COMES OUT OF THE BUS BUDGET RATHER THAN OUT OF
TASTE (new, session 27, FINDINGS 59.7). The frame affords 6.74 clocks a
byte after the measured decode and the audio, and a dual-address byte costs
9. So if B3 comes back saying the card cannot drive #EXREQ, the container
has to reach 27,995 B a frame — 328 KB/s of payload to fit at 12 fps, where
the gate container sits at 37,403 B and 438 KB/s: 34% too big.
Three things make this less alarming than the number looks, and one makes it worse:
- The gate container is deliberately the heaviest thing the encoder emits (span-heavy, the 488 recipe, every block mode exercised). It is a test fixture, not a shipping target.
- A lighter container also decodes cheaper, so the 68.5% decode term falls with the byte term. 328 KB/s is the pessimistic reading of the lever.
rc_fr_singe_scsi_cpufit.dlxalready exists — the encoder has had a CPU-fitting mode since session 11.- Worse:
15_bus_occupancy.pyREFUSES to price the cpufit container, correctly, because the C68K measurement it cross-checks against belongs to the gate container. So E7 starts with a harness re-run (tools/bench/c68k/run.sh) against whichever container is to be the target, and until that is done "34% too big" is a statement about the fixture rather than about the project.
G1. Import the scene graph — early, because it is a measurement input.
DONE, session 24 — FINDINGS 56. It was pulled ahead for exactly the reason
given, and it paid: the worst gap between two consecutive decision points is
zero, and 5.4% of the game's 612 branch transitions are. Two seeks can fall
back to back with no play between them, so 51.2's slack rule can be answered NO
by the content rather than by the buffer.
It does not break the design — a branch on an empty ring costs the prefill (149.7 ms, 1.80 frame slots at 488 KB/s), not the climb — but it removes the margin: at 488 KB/s in a 256 KB ring, 76% of this game's branch points arrive before the ring has refilled, and a 512 KB ring makes that 90%. The ring is not the lever; the surplus is.
Two constraints on the input layer came with it: the arcade needs eight directions, and the shortest input window is 98 ms against a 72.13/90.16 ms frame slot, so input cannot be polled on the frame tick (56.7).
The cross-check plan was wrong and is struck. The SNES chapters are
derived from DirkSimple, by their own README, so there is one transcription and
not two; the diff catches conversion errors only (56.2). Nothing is vendored:
tools/import/scenegraph.py is the one file coupled to those projects and it
writes this project's own DLXSCENE1 schema into gitignored tmp/
(USER DECISION, session 24).
M4 — the whole game
Listed for completeness; past M3 these are scope, not risk.
- C1. Full-disc survey, 22.8 minutes. Classify content / menu / bonus —
not menu vs content: the two largest streams are bonus material and look like
content by size, duration and bitrate alike (25.1). Run
07_motion_survey.pyper stream first for a hot-window shortlist. Gated by E4. - E4.
H.buildk-means, 51 s of a 55 s run, once per scene. The thing to attack before C1, and not anything in the per-frame path (27.6). - E2.
--spans allas default. (re-encode bundle item 2; see P2.) Still a recommendation, not a measurement (43.6.1), and the only loaded lever on the encoder's byte side (44.3). It spends every profitable byte, which raiseswire, which shrinkspipe - wire, which lengthens the refill climb after every branch. That interaction is not priced, and M3 is where it becomes measurable. - E3. Re-derive span selection jointly with
lam(39.3). (bundle item 3.) - C2. Framing — crop vs squash vs wide (FINDINGS 12). Needs an eyeball against arcade reference, not a measurement. Cheap; blocks only final encodes.
- C3. Disk image packaging, ~1.09 GiB at the candidate rate.
- G2/G3. Branching, input windows, death clips, attract mode; playtest.
Dependency summary
P4a DONE (59): the channel drives the data phase and
holds the bus -- 391.7% of a frame becomes 40..95%
│
B3 #EXREQ? ──┬─ YES ─> single address, 5 clk/B, 92.2% ── FITS ──┐
│ │
└─ NO ──> auto-request, 9 clk/B FLOOR, 110.2% ──> E7 byte target
(the frame affords 6.74; 59.7) 328 KB/s
│
P2 re-encode bundle (sector alignment is now a PRECONDITION, 59.4) ┤
P1 P3 P4b P5 P7, P6 bus cost (52), G1 scene graph (56) ────────────┼─> M2 ─>
B1 seek+rate (sets HEADROOM, not fit) ─────────────────────────────┘ M3 ─> M4
B2 blanking ─> (page 1; do not pre-build on it)
Read that top-left branch as the project's live question. Everything else
on the diagram is work; #EXREQ is a fact about a board nobody here has, and it
decides which of the two lower paths the player is on.
Standing rules that apply to all of it
- Green light first and last.
./tools/bench/check.sh, ALL GREEN, before and after. Never two MAME jobs at once — session 18 did it, twodecode.luaruns shared a log file, and it produced a 0-byte log and 15 wasted minutes. - Name the layer. Emulated, or real hardware. Every progress claim.
- Label measured / estimated / folklore. A rate with no provenance is folklore even when it is plausible, and this project has already paid for that twice.
- No new default constants. Rates stay explicit arguments. If a measurement is not available, report the sensitivity across several rates rather than picking one.