Pace the ring, then read the DMAC config out of the IPL ROM: audio is cheap and the disk is not
Two sessions that were never separated in the working tree, so they land as one commit. check.sh ALL GREEN before and after both. SESSION 19 -- the ring rig gets a frame clock (FINDINGS 51). src/player/stream.s had no frame clock: it asked for record i the instant it finished i-1, outran any finite pipe, and never let the ring back up. The 49.1 sweep passing at 48 KB was therefore a wrap-correctness result and nothing else. PACE/PACEON ($18034/$18038) hold the decoder to 12 fps, so FR_HEAD-FR_TAIL finally means what it reads as: whole frames the decoder could still draw with delivery stopped dead. PACEON=0 free-runs and is what the wrap gate still uses, so every figure in 49 is unmoved. Paced, on the gate container: 64 KB holds 2 frames, 256 KB holds 7-8, 512 KB holds 14-15, all pixel-exact. Tolerance is ceiling-1, measured by cutting the pipe: 256 KB buys 500 ms of dead pipe, not 583. SLACK IS ACCUMULATED, NOT OWNED. It is built out of pipe-wire and a seek spends all of it. At 488 KB/s a 256 KB ring needs 4.83 s of play to reach its ceiling from empty; 512 KB needs 8.42 s to reach 14. A bigger ring raises the ceiling AND lengthens the climb, so a branch point does not ask "is the buffer big enough" but "has there been enough play since the last one" -- and Dragon's Lair's decision points are seconds apart. The rig now also says WHICH resource is binding: at 460 KB/s every ring from 192 KB to 512 KB is rate-bound at ceiling 4 and never fills, so larger rings are dead RAM in that scene. 20_seek_slack.py is the same model rewritten in Python from record sizes, sharing no code with the Lua producer: 35/35 ceilings inside its bracket. SESSION 20 -- the DMAC configuration was in the IPL ROM the whole time (FINDINGS 52). ROADMAP's "do this first" was to put the ADPCM stream on the bus. That needs a clocks-per-byte figure for the audio channel, and 11_cpu_budget.py was charging audio the DISK's rate -- 5 clk/B, its own help text calling it "single-address, bus held". Audio was being charged the favourable end of B3, a 242 KB/s open question. It never had to be a guess. The IPL ROM programs all four HD63450 channels itself and MAME boots the rig with it, so 21_iplrom_dmac.py reads the configuration out of the image and decodes the MC68450 fields. Eight (address, expected bytes, meaning) sites; a mismatch or an unknown revision exits non-zero. In check.sh, no emulator, milliseconds. ch3 DCR=$80, OCR=$32: dual address, 8-bit port, cycle steal WITHOUT hold, REQG=10 external request. The DMAC arbitrates once per byte with no burst to amortise the 5..8 + 2 over, so an audio byte is 16..19 clocks, not 5 -- the old debit was 3.2x..3.8x small. And on the bus it is still nothing: 651 B/frame is 1.25%..1.48% of a frame, about 4% of what the decoder leaves. P6's bus risk does not materialise. The unit worry was worth checking and nearly right: 15.6 kHz is 8 MHz/512 = 15,625 samples/s, two 4-bit samples to a byte = 7,812.5 B/s exactly, and AUDIO_KBPS=7.8 is that in decimal kB while the tool multiplied by 1024. THE DISK CHANNEL IS PROGRAMMED IDENTICALLY. ch1 (SASI) is DCR=$80 too, and so is ch0. That is 16..19 clocks per delivered byte, where 42.4 brackets W at 5..12 and 42.5 has W=8 already missing 47/120 frames. The only worked example of a disk DMA configuration on this machine sits above the entire bracket, and at that price nothing fits at any container size. It is not scsiexrom.bin so B3 stays open -- what changed is that a cheap configuration is now the thing that has to be SHOWN. W <= 12 is a requirement on the player's DMAC programming, not a range the hardware hands us, and it is now the largest open number in the project, ahead of the rate. An unforced cross-check fell out: 15_bus_occupancy.py's new W sweep puts W=8 at 105.7% of the frame, agreeing with 42.5's 47/120, from mode histograms and bus clocks respectively, two models sharing no code. Also: ADPCM outranks the disk at the arbiter (CPR 1 against 2), so an audio byte never waits and a video byte does -- relevant to 51's smooth-rate delivery model. README MEDIA. stream.lua gains DLX_SNAP_EVERY=1 (needs DLX_PACE, off by default, on no path check.sh takes) and tools/media/make_readme_media.py turns the PNGs into docs/img/. The stills and both clips are MAME's own screen pixels. Building it turned up something worth recording. 116 of 119 captured frames are pixel-exact against dlx.py; three are TORN -- frame n on top, frame n-1 below the tear line -- because MAME captured the screen while the block loop was partway down it. decode.s writes straight to the displayed page (one display path, 28.1), so a real player tears the same way, and this is the first time that consequence has been visible rather than argued. The script ASSERTS the tear and refuses to build otherwise, rather than trimming three frames and reporting "every frame I kept is exact". Second correction the capture forced: the snapshot fires before frame n is decoded, so the obvious reading is that it holds frame n-1 -- it does not, because MAME renders the screen at the end of the machine frame, by which time the 68000 has finished frame n. 11_cpu_budget.py's "validated to within 1 pt" line is also corrected: the model reads 2..10 pt HIGH and by more as the frame gets harder, which was already true before either session. src/player/decode.s is unchanged; decode.bin is still 1,296 B at the same MD5. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
+142
-8
@@ -10,6 +10,25 @@ cd "$(dirname "$0")/../.."
|
||||
python3 tools/encoder/extract.py 00020 tmp/fr_00020 12 crop
|
||||
mkdir -p tmp/snap_verify tmp/snap256
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# THE ONE PLACE THE RETIRED PIPE FIGURE STILL LIVES. Session 18 removed it as
|
||||
# a default from every analysis tool and from tools/bench/stream.lua, because it
|
||||
# was never a bus measurement -- a user-supplied "4 Mbps" with no provenance,
|
||||
# 10% of SCSI-1's asynchronous rating (FINDINGS 42.1) -- and a default let table
|
||||
# after table be scored against it without anyone restating what it was.
|
||||
#
|
||||
# It survives HERE and only here because the gate container was ENCODED with it,
|
||||
# and every per-block and span constant in FINDINGS 41/43/45/49 is fitted to that
|
||||
# container. Changing this number is not an edit, it is a re-encode plus a
|
||||
# re-measurement of all of them.
|
||||
#
|
||||
# It is a CONTAINER RECIPE, not a claim about any medium. Do not read a delivery
|
||||
# rate out of it, do not copy it into a tool, and do not add a default anywhere
|
||||
# that would resurrect it. When the pipe is finally measured, this becomes an
|
||||
# ordinary encoder setting and the comment goes.
|
||||
GATE_SPAN_KBPS=488
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
run() { # run <script> <snapdir>
|
||||
rm -f "tmp/$2/x68000"/*.png
|
||||
( cd tmp && SDL_VIDEODRIVER=dummy timeout -k 5 120 mame x68000 -bios ipl10 \
|
||||
@@ -46,7 +65,10 @@ echo "--- session 12: the DLX3 span container round-trips (FINDINGS 41) ---"
|
||||
# header and is painted by the span section instead -- so this encodes, WRITES
|
||||
# the container, reads it back with the reference decoder and compares. It also
|
||||
# asserts that it emitted enough spans to have tested anything.
|
||||
python3 tools/analysis/16_span_roundtrip.py > tmp/span_roundtrip.log 2>&1 \
|
||||
# --kbps is required now (session 18): the tool has no default rate, so the gate
|
||||
# has to say which one it is testing at. Same recipe constant as the container.
|
||||
python3 tools/analysis/16_span_roundtrip.py --kbps $GATE_SPAN_KBPS \
|
||||
> tmp/span_roundtrip.log 2>&1 \
|
||||
|| { cat tmp/span_roundtrip.log; exit 1; }
|
||||
tail -4 tmp/span_roundtrip.log
|
||||
|
||||
@@ -69,16 +91,31 @@ echo "--- session 7: 68000 decoder is pixel-exact (FINDINGS 28) ---"
|
||||
# right if all 120 were.
|
||||
# The gate container is the HEAVIEST stream the encoder emits: the scsi mode
|
||||
# decision (the only profile left after session 9 dropped sasi on capacity,
|
||||
# FINDINGS 32) with the span pass drawing on the full 488 KB/s pipe, so every
|
||||
# frame carries a span table and all four block modes are still exercised.
|
||||
# FINDINGS 32) with the span pass drawing on a byte ceiling wide enough that
|
||||
# every frame carries a span table and all four block modes are still exercised.
|
||||
# That ceiling is GATE_SPAN_KBPS above -- a recipe, not a delivery rate.
|
||||
# Spans are the newest and least-proven path in decode.s; gating on a container
|
||||
# where they are rare would be gating on the old decoder. FINDINGS 41.
|
||||
DLX=tmp/rc_fr_singe_scsi_span.dlx
|
||||
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile scsi \
|
||||
--kbps 280 --span-kbps 488 --spans all
|
||||
python3 tools/bench/prep_dlx.py "$DLX" > tmp/prep_dlx.log
|
||||
# The rig loads the whole stream into a 2 MB machine, so a scsi window does not
|
||||
# fit and prep_dlx truncates it. Verify against exactly the prefix it emitted.
|
||||
--kbps 280 --span-kbps $GATE_SPAN_KBPS --spans all
|
||||
# RIG_RAM is the EMULATED MACHINE's memory, and it is not a claim about the
|
||||
# target. The rig preloads the whole container into RAM at 0x30000; the shipping
|
||||
# player streams from disk into a ring buffer and never holds a window at once,
|
||||
# so preloading is unlike the player at ANY size. At the 2 MB of a stock machine
|
||||
# this gate covered 37 of 120 frames (FINDINGS 44.6.4) -- the span-heavy
|
||||
# container is 5,261,814 B of stream, ending at 0x534BF6. 6 MB covers all 120.
|
||||
#
|
||||
# Raising it is licensed by measurement, not by convenience: at 2M and 6M the
|
||||
# five synthetic anchors come out BIT-IDENTICAL (40,729 / 921,187 / 1,376,881 /
|
||||
# 1,229,883 / 506,533 cycles) despite sitting at different addresses in the two
|
||||
# layouts, so MAME's cycle model does not depend on ramsize over this range.
|
||||
# FINDINGS 45. What is still NOT tested, at either size, is the streaming path.
|
||||
RIG_RAM=${RIG_RAM:-6}
|
||||
python3 tools/bench/prep_dlx.py "$DLX" --ram $((RIG_RAM * 0x100000)) > tmp/prep_dlx.log
|
||||
# Verify against exactly the frame list prep_dlx emitted. It no longer truncates
|
||||
# at the default RIG_RAM, but the guard stays: lower RIG_RAM, or a heavier
|
||||
# container, brings truncation straight back and it must stay announced.
|
||||
NF=$(sed -n 's/.*nframes=\([0-9]*\),.*/\1/p' tmp/decode_meta.lua)
|
||||
grep -a "TRUNCATED" tmp/prep_dlx.log || true
|
||||
tools/vasm/vasmm68k_mot -Fbin -o tmp/decode.bin src/player/decode.s > /dev/null
|
||||
@@ -93,7 +130,7 @@ rm -f tmp/snap_decode/x68000/*.png
|
||||
# compared a partially drawn screen and reported 49,005 differing pixels, which
|
||||
# reads as a decoder bug and is not one.
|
||||
( cd tmp && DLX_VERIFY_ONLY=1 SDL_VIDEODRIVER=dummy stdbuf -oL timeout -k 5 300 mame x68000 \
|
||||
-bios ipl10 -ramsize 2M -video soft -window -sound none -nothrottle -plugins \
|
||||
-bios ipl10 -ramsize ${RIG_RAM}M -video soft -window -sound none -nothrottle -plugins \
|
||||
-autoboot_script ../tools/bench/decode.lua \
|
||||
-snapshot_directory ./snap_decode -snapview native -seconds_to_run 60 \
|
||||
> decode_check.log 2>&1 )
|
||||
@@ -130,4 +167,101 @@ else
|
||||
echo " SKIPPED: no px68k at $PX68K (set PX68K= to point at a checkout)"
|
||||
fi
|
||||
|
||||
echo "--- session 20: the DMAC config, read out of the IPL ROM (FINDINGS 52) ---"
|
||||
# The audio and disk per-byte debits are no longer a recollection about the
|
||||
# HD63450: they are bytes at named addresses in the ROM MAME boots this rig
|
||||
# with. This gate re-reads them. It is cheap, it needs no emulator, and if a
|
||||
# different ROM revision is ever pointed at it, it says so rather than decoding
|
||||
# some other code and reporting a number.
|
||||
# Skipped rather than failed when the ROM is not where MAME keeps it: that is a
|
||||
# path outside this repo.
|
||||
IPLROM=${IPLROM:-$HOME/mame/roms/iplrom.dat}
|
||||
if [ -f "$IPLROM" ]; then
|
||||
python3 tools/analysis/21_iplrom_dmac.py "$IPLROM" > tmp/iplrom_dmac.log 2>&1 \
|
||||
|| { cat tmp/iplrom_dmac.log; exit 1; }
|
||||
grep -ac "^ OK " tmp/iplrom_dmac.log | xargs printf " %s evidence sites hold; "
|
||||
sed -n 's/^ = \(.*clocks per audio byte\)/audio is \1/p' tmp/iplrom_dmac.log
|
||||
else
|
||||
echo " SKIPPED: no IPL ROM at $IPLROM (set IPLROM= to point at it)"
|
||||
fi
|
||||
|
||||
echo "--- session 18: the shared-body split is a no-op (FINDINGS 49.7.5) ---"
|
||||
# src/player/decode.s and src/player/stream.s assemble from ONE copy of the block
|
||||
# loop and the span chain (src/player/frame.i) so that the two front-ends cannot
|
||||
# drift apart. The drift would be silent -- both would still decode correctly,
|
||||
# and only the cost model would be wrong, because the 66.0 clocks/span, 9.143
|
||||
# clocks/coarse pixel and every per-block constant in FINDINGS 24/30/40/41 are
|
||||
# fitted to those exact bytes. So the split is asserted to be a no-op rather than
|
||||
# assumed to be one.
|
||||
DECODE_MD5=7a7a06f8c6d097ee0041bca4aefa3eb2 # decode.bin before the split, 1296 B
|
||||
GOT=$(md5sum tmp/decode.bin | cut -d" " -f1)
|
||||
[ "$GOT" = "$DECODE_MD5" ] || {
|
||||
echo "FAIL: decode.bin is $GOT, expected $DECODE_MD5 ($(stat -c%s tmp/decode.bin) B)."
|
||||
echo " The block loop or the span chain changed. That is allowed -- but"
|
||||
echo " every cycle constant in FINDINGS 24/30/40/41 is fitted to the old"
|
||||
echo " bytes, so re-measure them and move this hash, do not just move it."
|
||||
exit 1; }
|
||||
echo " decode.bin unchanged at $(stat -c%s tmp/decode.bin) B ($DECODE_MD5)"
|
||||
# Same argument for the loader maths, which prep_dlx.py and prep_stream.py now
|
||||
# share via tools/bench/dlxload.py: a second copy of the palette packing would
|
||||
# drift and the symptom would be wrong colours in one rig only.
|
||||
python3 tools/bench/prep_dlx.py "$DLX" --ram $((RIG_RAM * 0x100000)) --out tmp/_pdchk > /dev/null
|
||||
cmp -s tmp/_pdchk_data.bin tmp/decode_data.bin || {
|
||||
echo "FAIL: prep_dlx.py is not reproducible"; exit 1; }
|
||||
echo " prep_dlx.py blob reproducible ($(stat -c%s tmp/decode_data.bin) B)"
|
||||
rm -f tmp/_pdchk_data.bin tmp/_pdchk_meta.lua
|
||||
|
||||
echo "--- session 18: 120 frames through a bounded RING (FINDINGS 49) ---"
|
||||
# The gate above preloads the whole container into RAM and proves the DECODER.
|
||||
# This proves the DELIVERY path: the same 120 frames decoded out of a 256 KB
|
||||
# ring on a STOCK 2 MB machine, with the container in a host file. The block
|
||||
# loop reads with a monotonically increasing a0 and no bounds check, so a record
|
||||
# placed wrongly by the wrap policy corrupts pixels rather than faulting -- which
|
||||
# is why this is gated on the same pixel-exact comparison and not on a checksum.
|
||||
tools/vasm/vasmm68k_mot -Fbin -o tmp/stream.bin src/player/stream.s > /dev/null
|
||||
python3 tools/bench/prep_stream.py "$DLX" > tmp/prep_stream.log
|
||||
mkdir -p tmp/snap_stream
|
||||
rm -f tmp/snap_stream/x68000/*.png
|
||||
( cd tmp && DLX_STREAM_KBPS=0 SDL_VIDEODRIVER=dummy stdbuf -oL timeout -k 5 600 \
|
||||
mame x68000 -bios ipl10 -ramsize 2M -video soft -window -sound none \
|
||||
-nothrottle -plugins -autoboot_script ../tools/bench/stream.lua \
|
||||
-snapshot_directory ./snap_stream -snapview native -seconds_to_run 90 \
|
||||
> stream_check.log 2>&1 )
|
||||
# Same truncation trap as the decode stage: without this, a run that exited
|
||||
# mid-decode is compared against a half-drawn screen and reads as a wrap bug.
|
||||
grep -q "snapshot taken" tmp/stream_check.log || {
|
||||
echo "FAIL: the ring-buffer pass did not complete -- no snapshot marker."
|
||||
tail -5 tmp/stream_check.log; exit 1; }
|
||||
grep -a "ring: \|DEADLINE" tmp/stream_check.log | sed "s/\[STR\] / /"
|
||||
python3 tools/bench/verify_decode.py "$DLX" --snap tmp/snap_stream
|
||||
|
||||
echo "--- session 19: the PACED ring, and what a branch point costs (FINDINGS 51) ---"
|
||||
# The stage above runs the ring FREE-RUNNING, which is right for what it gates:
|
||||
# an unlimited pipe removes delivery as a variable and leaves the wrap policy
|
||||
# alone under test. It cannot see buffering, because a decoder that never waits
|
||||
# never lets the ring back up -- 49.7.2, and it is why 48 KB passed while
|
||||
# holding one record. This runs the same 120 frames with the decoder held to
|
||||
# 12 fps, which is the only configuration in which FR_HEAD-FR_TAIL means what
|
||||
# it is read to mean.
|
||||
#
|
||||
# Gated on: pixel-exact, zero UNDERRUNS, and a ceiling that has not moved. The
|
||||
# ceiling is a property of THIS container in a 256 KB ring; it is asserted
|
||||
# rather than printed because a change in it is a change in how much a branch
|
||||
# point can afford, and that should not slip through as a line in a log.
|
||||
bash tools/bench/pace_run.sh 256 0 > tmp/pace_check.log 2>&1 || {
|
||||
echo "FAIL: the paced ring pass did not complete."; tail -8 tmp/pace_check.log
|
||||
exit 1; }
|
||||
grep -aE "SEEK SLACK|UNDERRUNS" tmp/pace_check.log
|
||||
grep -q "UNDERRUNS: 0/120" tmp/pace_check.log || {
|
||||
echo "FAIL: the paced decoder underran -- a frame's slot arrived before its"
|
||||
echo " record did. Free-running this is earliness (49.6); paced it is not."
|
||||
exit 1; }
|
||||
grep -q "ceiling 8 frames" tmp/pace_check.log || {
|
||||
echo "FAIL: the 256 KB seek-slack ceiling is no longer 8 frames (FINDINGS 51)."
|
||||
echo " Re-run tools/bench/pace_sweep.sh and re-derive 51 before editing"
|
||||
echo " this number -- it is what a branch point can spend."
|
||||
exit 1; }
|
||||
grep -q "^OK" tmp/pace_check.log || { echo "FAIL: paced pass not pixel-exact";
|
||||
tail -4 tmp/pace_check.log; exit 1; }
|
||||
|
||||
echo "ALL GREEN"
|
||||
|
||||
Reference in New Issue
Block a user