11_cpu_budget.py takes --machine. Clocks confirmed from MAME 0.277
x68k.cpp:1133/1194/1200, not recalled: x68000 AND x68ksupr are both
40_MHz_XTAL/4 = 10 MHz; only the XVI is faster at 33.33_MHz_XTAL/2.
sasi scsi
stock 10MHz 31% miss 42% miss
XVI 16.7MHz 0% miss 0% miss
sasi is the cheaper profile but it does not fit either at 10 MHz. The XVI
column is headroom, not a target: the profiles are an I/O-bandwidth axis and
say nothing about CPU, and the locked target CPU is a stock 10 MHz 68000 for
both of them. So both profiles have to fit the same 833,333-cycle budget, and
the cycle ceiling has to be enforced in the encoder regardless of which one
ships.
Model comparisons are now gated to the clock and framerate they were stated
at: quoting 24.5's 76.6% or the stock-machine 68000 timings against an XVI
budget compares a model to a measurement of a different machine.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
src/player/decode.s parses DLX1 and decodes straight into GVRAM. Verified
pixel-exact over a 120-frame sequential run of the worst sustained window on
the disc -- all four block modes, full temporal recursion, so the last frame
is only right if all 120 were. In check.sh.
It costs a mean of 81.7% of a 12fps frame budget, and 31% of frames exceed
100% (42% at scsi). CPU is now the binding constraint. FINDINGS 28.
Three things that were believed and are not true:
- The dual-display-path plan of FINDINGS 24.5/25.6 is incoherent. The compose
path needs a RAM copy of the previous reconstruction; the direct path's
selling point is that it keeps none. Mixing them shows stale pixels on 70 of
120 frames, worst frame 18.8% of the screen. Every coherent repair is dearer
than not mixing, and 24.5's two figures were both copies with no decode in
either, so there was never a crossover to find. One path ships, and the 96KB
reference frame is gone. tools/analysis/10_pathmix_drift.py keeps the
counterexample runnable; check.sh asserts it still reproduces.
- The four block modes do not cost the same. V1 300, V4 448, RAW 400 cycles
against the old model's flat 207.8. V4 is 25% of blocks and 50% of the
cycles, and the mode decision charges it bytes it does not charge cycles for.
tools/analysis/11_cpu_budget.py reproduces all four frames timed on the
68000 to within 1 point. Hand-derived timings agree to 0.5% on V1.
- The container is big-endian but not aligned. Variable-length records laid end
to end put frame 1's length field at an odd address, and move.l (a0)+ there
is an address error: frame 0 decoded perfectly and then vectored into the
IPL for 59 emulated seconds looking like a hang. Found by dumping PC, not by
reading the source.
Also: an all-V1 frame, the cheapest possible full redraw, is 110.5% of budget.
No mode assignment fits a scene cut at 12fps. That one needs a decision, not a
measurement.
Next: charge cycles in the mode decision and bisect against 833,333 per frame,
the way session 6 bisects lam against bytes -- but with no bucket, because a
late frame cannot be banked.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6