Three things, and the last one reversed itself when the datasheet arrived.
A SECOND EMULATOR. tools/bench/c68k/ links px68k's C68K core into a headless
harness -- no SDL, no ROMs, no emulated machine, because the decoder touches
nothing but RAM, the control block and GVRAM. decode.s is now pixel-exact under
two independent CPU cores, and cycle-table error against MAME is bounded at
3.3%, running against us. MAME 0.277's M68000 turns out to be the MICROCODE
core, not Musashi (m68000.lst + m68000gen.py), so this is two structurally
different timing models agreeing rather than two tables. FINDINGS 28.8's "V4
costs more than RAW" reproduces independently. FINDINGS 37.
THE BUS. Nothing since FINDINGS 24 had counted the 68000's local memory bus --
one 4-clock cycle at a time, carrying instruction prefetch as well as data. The
decoder occupies 86.7% of it and PREFETCH IS 62% OF THAT TRAFFIC, so a data-only
count understates occupancy by 2x. Two sources check each other: c68k_bench
counts every bus callback exactly, and a static walk of decode.lst supplies the
prefetch no emulator here can report. The walk reproduces the measured data half
to 0.04%, which is what licenses its prefetch half, and 15_bus_occupancy.py is a
gate rather than a report because every bus figure depends on that check.
FINDINGS 38.
THE DMAC CHAIN LOSES. FINDINGS 29.6 named it the one uncosted lever. Costed from
bus arithmetic -- a read cycle plus a write cycle, 8 clocks a pixel -- it scored
1/120 frames over budget against the v6 span's 10/120 and looked decisive. Then
the MC68450 manual (Motorola Jul 1989, now at ~/src/mc68450.pdf): Fig 4-25 sheet
4 puts a dual-address word between two 16-bit ports at 9 CLOCKS, because note 2
gives the DMAC 4-clock reads and 5-clock WRITES. The 68000 writes in 4.
DMAC 9.000 clocks/pixel datasheet
v6 9.152 clocks/pixel measured, FINDINGS 30
1.7%. Scored additively, 86% of what remains of the DMAC's advantage is v6's
24-pixel padding quantum -- a property of its unrolled movem chain, fixable in
software with a finer tail chain, worth 55/120 -> 18/120 against the DMAC's
12/120. Recommendation: fix the quantum, drop the DMAC. Six frames does not buy
a reserved channel, a two-region container layout and a timing dependency
neither emulator here can verify. The container is identical either way -- v6's
record and an HD63450 chaining entry are both 6 bytes, so the chain array IS the
span table -- so nothing is foreclosed. FINDINGS 39.
TWO CORRECTIONS TO MY OWN WORK IN THE SAME SESSION:
- I argued FINDINGS 35's flat CPU debit for the disk was too pessimistic and
rescored the window at 53/120 with max(CPU, bus). Wrong. A 68000 has no cache
and a two-word prefetch queue, so it stalls the moment another master takes
the bus, and the MC68450 hands the bus over in SLABS under limited-rate
auto-request rather than interleaving per operand. DMA is additive. 84/120
stands and 14_dmac_chain.py reproduces it exactly. What 86.7% occupancy really
says is that there is almost no room to overlap anything. FINDINGS 38.3.
- The first DMAC costing was derived where a primary source existed. Both wrong
answers were confident and both were caught by reading the manual.
Also landed:
- FINDINGS 5's 8 clocks/word for the SCSI DMA, STATUS's own "most load-bearing
unmeasured number", is now bracketed by the datasheet: 5 clk/word with the bus
held, ~12 if the DMAC arbitrates per word. 8 is a supported midpoint, and
which end applies is a player design decision worth 7 clocks a word on a
480 KB/s stream. FINDINGS 39.7.
- check.sh gains two gates: the C68K pixel-exact decode (seconds, no MAME) and
the bus-model self-check. Both skip cleanly without a px68k checkout.
- spanned blocks are now charged their mode-map dispatch, which FINDINGS 30.7
flagged as uncounted in 12_span_tradeoff.py.
- MAME timed runs must be budgeted by WALL CLOCK, not -seconds_to_run: this box
runs x68000 at ~0.033x realtime and two runs were killed by their own timeout.
That is why the all-RAW cell in 37.3 is empty. The C68K harness does the same
work in seconds because it emulates a CPU and not a machine.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
122 lines
6.5 KiB
Bash
Executable File
122 lines
6.5 KiB
Bash
Executable File
#!/bin/bash
|
|
# Green-light check: re-runs both display regression tests AND the rate-control
|
|
# drift test, from the Blu-ray. ~2 min. Run from the repo root. Any non-zero
|
|
# exit means something drifted.
|
|
set -e
|
|
cd "$(dirname "$0")/../.."
|
|
[ -d /media/reala-misaki/BDROM ] || {
|
|
echo "Blu-ray not mounted. udisksctl loop-setup -r -f DRAGONS_LAIR.iso"; exit 2; }
|
|
|
|
python3 tools/encoder/extract.py 00020 tmp/fr_00020 12 crop
|
|
mkdir -p tmp/snap_verify tmp/snap256
|
|
|
|
run() { # run <script> <snapdir>
|
|
rm -f "tmp/$2/x68000"/*.png
|
|
( cd tmp && SDL_VIDEODRIVER=dummy timeout -k 5 120 mame x68000 -bios ipl10 \
|
|
-video soft -window -sound none -nothrottle -plugins \
|
|
-autoboot_script "../tools/bench/$1" \
|
|
-snapshot_directory "./$2" -snapview native -seconds_to_run 6 >"$2.log" 2>&1 )
|
|
}
|
|
|
|
echo "--- session 3: 768-wide IPL timing (FINDINGS 22) ---"
|
|
python3 tools/bench/prep_frame.py tmp/fr_00020 tmp/frame.bin 0
|
|
run show_frame.lua snap_verify
|
|
python3 tools/bench/verify_frame.py
|
|
|
|
echo "--- session 4: real 256x256 mode (FINDINGS 23) ---"
|
|
python3 tools/bench/prep_frame.py tmp/fr_00020 tmp/frame256.bin 0 --reserve-black
|
|
run show_frame256.lua snap256
|
|
python3 tools/bench/verify_frame256.py
|
|
|
|
echo "--- session 6: rate-control drift (FINDINGS 26/27) ---"
|
|
# The codec is temporally recursive, so a rate controller can report quality for
|
|
# a reconstruction no decoder will ever produce -- silently. This asserts that a
|
|
# decoder replaying the emitted stream rebuilds exactly what the encoder
|
|
# recorded. ~55 s, nearly all of it k-means in H.build.
|
|
[ -d tmp/fr_singe ] || python3 tools/encoder/extract.py 00223 tmp/fr_singe 12 crop 539.4 10.0
|
|
# NOT piped into tail: a pipeline's exit status is the last command's, which
|
|
# would swallow the failure this whole script exists to catch.
|
|
python3 tools/analysis/09_ratectl_drift.py > tmp/drift_check.log 2>&1 \
|
|
|| { cat tmp/drift_check.log; exit 1; }
|
|
tail -9 tmp/drift_check.log
|
|
|
|
echo "--- session 7: display-path coherency (FINDINGS 28.1) ---"
|
|
# 10_pathmix_drift.py is a COUNTEREXAMPLE, kept runnable: the dual-path plan of
|
|
# FINDINGS 24.5/25.6 must still be shown to corrupt frames, and the strategy the
|
|
# player actually uses must still be clean. A green light here means the reason
|
|
# decode.s has one display path is still demonstrable, not just asserted.
|
|
python3 tools/analysis/10_pathmix_drift.py > tmp/pathmix.log 2>&1 \
|
|
&& { echo "FAIL: the dual-path plan no longer reproduces its own defect"; \
|
|
cat tmp/pathmix.log; exit 1; }
|
|
grep -a "frames displaying pixels" tmp/pathmix.log
|
|
python3 tools/analysis/10_pathmix_drift.py --fix direct > tmp/pathmix_direct.log 2>&1 \
|
|
|| { echo "FAIL: direct-to-GVRAM is no longer coherent"; cat tmp/pathmix_direct.log; exit 1; }
|
|
|
|
echo "--- session 7: 68000 decoder is pixel-exact (FINDINGS 28) ---"
|
|
# The strongest display test in the tree: 120 frames decoded in sequence by
|
|
# 68000 code, every block mode, full temporal recursion. A SKIP block is a claim
|
|
# about the previous frame still being on screen, so the last frame is only
|
|
# right if all 120 were.
|
|
# The gate container is the CURRENT default encode: scsi (the only profile left
|
|
# after session 9 dropped sasi on capacity, FINDINGS 32), cost-aware mode
|
|
# decision on, DLX2 4-byte-aligned records. It is also the heavier stream --
|
|
# 43% RAW against sasi's 10% -- so it exercises the decoder harder than the
|
|
# session-7 container this gate used to run on.
|
|
DLX=tmp/rc_fr_singe_scsi_cpufit.dlx
|
|
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile scsi
|
|
python3 tools/bench/prep_dlx.py "$DLX" > tmp/prep_dlx.log
|
|
# The rig loads the whole stream into a 2 MB machine, so a scsi window does not
|
|
# fit and prep_dlx truncates it. Verify against exactly the prefix it emitted.
|
|
NF=$(sed -n 's/.*nframes=\([0-9]*\),.*/\1/p' tmp/decode_meta.lua)
|
|
grep -a "TRUNCATED" tmp/prep_dlx.log || true
|
|
tools/vasm/vasmm68k_mot -Fbin -o tmp/decode.bin src/player/decode.s > /dev/null
|
|
mkdir -p tmp/snap_decode
|
|
rm -f tmp/snap_decode/x68000/*.png
|
|
# stdbuf -oL: a FILE is block-buffered too, so without it a long MAME run is
|
|
# unobservable until it exits and a run that is merely finishing looks exactly
|
|
# like one that is wedged (FINDINGS 34.1).
|
|
# -seconds_to_run must cover the WHOLE sequential pass. The scsi container is
|
|
# 2.7x the payload of the session-7 one this gate used to run on, and at 20 s
|
|
# the pass was truncated -- MAME exited mid-decode and verify_decode.py then
|
|
# compared a partially drawn screen and reported 49,005 differing pixels, which
|
|
# reads as a decoder bug and is not one.
|
|
( cd tmp && DLX_VERIFY_ONLY=1 SDL_VIDEODRIVER=dummy stdbuf -oL timeout -k 5 300 mame x68000 \
|
|
-bios ipl10 -ramsize 2M -video soft -window -sound none -nothrottle -plugins \
|
|
-autoboot_script ../tools/bench/decode.lua \
|
|
-snapshot_directory ./snap_decode -snapview native -seconds_to_run 45 \
|
|
> decode_check.log 2>&1 )
|
|
# A truncated run must fail as a truncated run. Without this the only symptom is
|
|
# a pixel diff against a half-drawn frame.
|
|
grep -q "snapshot taken" tmp/decode_check.log || {
|
|
echo "FAIL: the 68000 sequential pass did not complete -- no snapshot marker."
|
|
echo " Raise -seconds_to_run; the pass needs the whole container decoded."
|
|
tail -5 tmp/decode_check.log; exit 1; }
|
|
python3 tools/bench/verify_decode.py "$DLX" --nframes "$NF"
|
|
|
|
echo "--- session 10: the same decode on a second CPU core (FINDINGS 37) ---"
|
|
# A SECOND emulator, and the cheapest strong test in the tree: seconds, no MAME,
|
|
# no ROMs. px68k's C68K core has its own cycle table and its own memory model,
|
|
# so a pass here says decode.s is pixel-exact under two independent cores and
|
|
# that the harness's byte-swapped RAM / high-byte-discarding GVRAM is right --
|
|
# which is what licenses its cycle and bus numbers.
|
|
# Skipped rather than failed when px68k is not checked out: it is an external
|
|
# tree, not part of this repo.
|
|
PX68K=${PX68K:-$HOME/src/px68k}
|
|
if [ -f "$PX68K/m68000/c68k.c" ]; then
|
|
make -s -C tools/bench/c68k PX68K="$PX68K"
|
|
bash tools/bench/c68k/run.sh tmp/c68k_frames.csv 2>tmp/c68k.log
|
|
grep -a "sequential pass" tmp/c68k.log
|
|
python3 tools/bench/c68k/verify_c68k.py "$DLX" --nframes "$NF"
|
|
|
|
echo "--- session 10: the bus model still matches the machine (FINDINGS 38) ---"
|
|
# 15_bus_occupancy.py derives instruction prefetch, which no emulator here can
|
|
# report, and validates itself against the DATA accesses the harness counts.
|
|
# If that check ever stops holding, every bus figure in FINDINGS 38/39 is
|
|
# unfounded -- so it is a gate, not a report.
|
|
python3 tools/analysis/15_bus_occupancy.py "$DLX" | sed -n '3,7p'
|
|
else
|
|
echo " SKIPPED: no px68k at $PX68K (set PX68K= to point at a checkout)"
|
|
fi
|
|
|
|
echo "ALL GREEN"
|