Build v7 into the player, and find the cost model 18% wrong on the block it made commonest
src/player/decode.s now paints v7 literal spans, pixel-exact under MAME and px68k's C68K core over a container where every frame carries 128-216 spans covering up to 38% of the picture. The span pass is blit.s v7 verbatim: the 66.0/9.143/9.978 fit was measured on that instruction sequence. The container is DLX3 -- a span section between the mode header and the block payload, since that is the only place the 68000 can reach without first parsing something of variable length. 16_span_roundtrip.py gates it in check.sh, and asserts it emitted enough spans to have tested anything. Two synthetic all-SPAN anchors price v7 inside decode.s at 151.2 and 225.6 clocks per 4x4 block, against FINDINGS 40's table of 151 and 226 -- 0.2% on both emulators. The measured mode costs what it was said to cost. Two things that were not on the list: TWO BYTE BUDGETS. FINDINGS 40's 18/120 was scored against the 488 KB/s PIPE, not the 280 KB/s profile, and at the profile rate the lam search has already spent the allowance -- spans fired on 5 frames of 120 and looked like a regression. The profile is a chosen quality rate point; the pipe is hardware. --kbps and --span-kbps are now separate and spans run before mu, because a span pays in bytes and mu pays in picture. Delivered: 86/120 over budget without spans, 77/120 at the profile budget, 34/120 on the pipe for +0.36 dB. C_SKIP_MIXED WAS NEVER MEASURED, and it was 18% low -- 45.0, now 55.0. It is the one constant in the table that came from a derivation, because the synthetic frame that would measure it cannot exist: a byte needs a coded block for its SKIP to be mixed. Four bracketing anchors measure it on both emulators with the header byte rotated through all four positions, and the partner mode solves back to its own anchored value to 0.2%. With it corrected the model predicts a real spanned decode to -0.06% mean / 0.09% worst, against -2.99% / 4.30%. It matters because a span marks its run SKIP, so mixed SKIPs dominate exactly the frames spans are judged on. Also: the rig had been writing its synthetic timing frames 26 KB past the top of a 2 MB machine, and got away with it because the modes it overran are data-independent. A span's jump displacements come out of the stream, so it is not. And frames-over-budget is no longer a safe headline -- the controller aims at the deadline, so 55 of 120 frames sit within 5% of it and a 1% cost shift moves 22 frames. FINDINGS 41. check.sh ALL GREEN, now gating on a span-heavy DLX3 container. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
+20
-8
@@ -40,6 +40,16 @@ python3 tools/analysis/09_ratectl_drift.py > tmp/drift_check.log 2>&1 \
|
||||
|| { cat tmp/drift_check.log; exit 1; }
|
||||
tail -9 tmp/drift_check.log
|
||||
|
||||
echo "--- session 12: the DLX3 span container round-trips (FINDINGS 41) ---"
|
||||
# 09 above replays SKIP semantics in Python and never reads a container. A v7
|
||||
# span breaks exactly that shortcut -- a spanned block reads SKIP in the mode
|
||||
# header and is painted by the span section instead -- so this encodes, WRITES
|
||||
# the container, reads it back with the reference decoder and compares. It also
|
||||
# asserts that it emitted enough spans to have tested anything.
|
||||
python3 tools/analysis/16_span_roundtrip.py > tmp/span_roundtrip.log 2>&1 \
|
||||
|| { cat tmp/span_roundtrip.log; exit 1; }
|
||||
tail -4 tmp/span_roundtrip.log
|
||||
|
||||
echo "--- session 7: display-path coherency (FINDINGS 28.1) ---"
|
||||
# 10_pathmix_drift.py is a COUNTEREXAMPLE, kept runnable: the dual-path plan of
|
||||
# FINDINGS 24.5/25.6 must still be shown to corrupt frames, and the strategy the
|
||||
@@ -57,13 +67,15 @@ echo "--- session 7: 68000 decoder is pixel-exact (FINDINGS 28) ---"
|
||||
# 68000 code, every block mode, full temporal recursion. A SKIP block is a claim
|
||||
# about the previous frame still being on screen, so the last frame is only
|
||||
# right if all 120 were.
|
||||
# The gate container is the CURRENT default encode: scsi (the only profile left
|
||||
# after session 9 dropped sasi on capacity, FINDINGS 32), cost-aware mode
|
||||
# decision on, DLX2 4-byte-aligned records. It is also the heavier stream --
|
||||
# 43% RAW against sasi's 10% -- so it exercises the decoder harder than the
|
||||
# session-7 container this gate used to run on.
|
||||
DLX=tmp/rc_fr_singe_scsi_cpufit.dlx
|
||||
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile scsi
|
||||
# The gate container is the HEAVIEST stream the encoder emits: the scsi mode
|
||||
# decision (the only profile left after session 9 dropped sasi on capacity,
|
||||
# FINDINGS 32) with the span pass drawing on the full 488 KB/s pipe, so every
|
||||
# frame carries a span table and all four block modes are still exercised.
|
||||
# Spans are the newest and least-proven path in decode.s; gating on a container
|
||||
# where they are rare would be gating on the old decoder. FINDINGS 41.
|
||||
DLX=tmp/rc_fr_singe_scsi_span.dlx
|
||||
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile scsi \
|
||||
--kbps 280 --span-kbps 488 --spans all
|
||||
python3 tools/bench/prep_dlx.py "$DLX" > tmp/prep_dlx.log
|
||||
# The rig loads the whole stream into a 2 MB machine, so a scsi window does not
|
||||
# fit and prep_dlx truncates it. Verify against exactly the prefix it emitted.
|
||||
@@ -83,7 +95,7 @@ rm -f tmp/snap_decode/x68000/*.png
|
||||
( cd tmp && DLX_VERIFY_ONLY=1 SDL_VIDEODRIVER=dummy stdbuf -oL timeout -k 5 300 mame x68000 \
|
||||
-bios ipl10 -ramsize 2M -video soft -window -sound none -nothrottle -plugins \
|
||||
-autoboot_script ../tools/bench/decode.lua \
|
||||
-snapshot_directory ./snap_decode -snapview native -seconds_to_run 45 \
|
||||
-snapshot_directory ./snap_decode -snapview native -seconds_to_run 60 \
|
||||
> decode_check.log 2>&1 )
|
||||
# A truncated run must fail as a truncated run. Without this the only symptom is
|
||||
# a pixel diff against a half-drawn frame.
|
||||
|
||||
Reference in New Issue
Block a user