The 68000 decoder draws pixel-exact frames, and does not fit
src/player/decode.s parses DLX1 and decodes straight into GVRAM. Verified pixel-exact over a 120-frame sequential run of the worst sustained window on the disc -- all four block modes, full temporal recursion, so the last frame is only right if all 120 were. In check.sh. It costs a mean of 81.7% of a 12fps frame budget, and 31% of frames exceed 100% (42% at scsi). CPU is now the binding constraint. FINDINGS 28. Three things that were believed and are not true: - The dual-display-path plan of FINDINGS 24.5/25.6 is incoherent. The compose path needs a RAM copy of the previous reconstruction; the direct path's selling point is that it keeps none. Mixing them shows stale pixels on 70 of 120 frames, worst frame 18.8% of the screen. Every coherent repair is dearer than not mixing, and 24.5's two figures were both copies with no decode in either, so there was never a crossover to find. One path ships, and the 96KB reference frame is gone. tools/analysis/10_pathmix_drift.py keeps the counterexample runnable; check.sh asserts it still reproduces. - The four block modes do not cost the same. V1 300, V4 448, RAW 400 cycles against the old model's flat 207.8. V4 is 25% of blocks and 50% of the cycles, and the mode decision charges it bytes it does not charge cycles for. tools/analysis/11_cpu_budget.py reproduces all four frames timed on the 68000 to within 1 point. Hand-derived timings agree to 0.5% on V1. - The container is big-endian but not aligned. Variable-length records laid end to end put frame 1's length field at an odd address, and move.l (a0)+ there is an address error: frame 0 decoded perfectly and then vectored into the IPL for 59 emulated seconds looking like a hang. Found by dumping PC, not by reading the source. Also: an all-V1 frame, the cheapest possible full redraw, is 110.5% of budget. No mode assignment fits a scene cut at 12fps. That one needs a decision, not a measurement. Next: charge cycles in the mode decision and bisect against 833,333 per frame, the way session 6 bisects lam against bytes -- but with no bucket, because a late frame cannot be banked. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -0,0 +1,136 @@
|
||||
#!/usr/bin/env python3
|
||||
"""REGRESSION TEST for the dual-display-path coherency defect (FINDINGS 28).
|
||||
|
||||
FAILS ON PURPOSE for the player design FINDINGS 24.5/25.6 specified, which is
|
||||
why it exists: it is the counterexample, kept runnable. Pass `--fix <strategy>`
|
||||
to check that a proposed player is coherent instead.
|
||||
|
||||
The defect. Two display paths were specified and priced, and the plan was to
|
||||
pick between them per frame on the non-SKIP block count:
|
||||
|
||||
compose-in-RAM-then-blit flat 53.6% of a 12fps frame
|
||||
decode-direct-to-GVRAM 76.6% x (non-SKIP fraction), and -- quoting
|
||||
FINDINGS 24.5 -- "no RAM reference frame is needed"
|
||||
|
||||
Both statements are true in isolation and incompatible together. The compose
|
||||
path's whole job is to assemble a FULL frame in RAM so the blit can be
|
||||
row-linear, and the pixels it does not decode this frame (the SKIP blocks) can
|
||||
only come from a RAM copy of the previous reconstruction. The direct path
|
||||
deliberately never writes that copy. So every direct frame silently invalidates
|
||||
the reference the next compose frame reads, and the stale pixels go to screen.
|
||||
|
||||
This is FINDINGS 26 again in different clothing: two code paths that disagree
|
||||
about what "the previous frame" means. 26 was caught between encoder and
|
||||
decoder; this one is between a decoder and itself.
|
||||
|
||||
Strategies (--fix):
|
||||
none the specified player: direct writes GVRAM only. UNSOUND
|
||||
dual direct also writes the RAM reference (costs ~1.52x). sound
|
||||
resync on direct->compose, re-read GVRAM into RAM first. sound
|
||||
compose never use the direct path. sound
|
||||
direct never use the compose path. sound
|
||||
|
||||
Needs a container: defaults to the worst sustained window on the disc at the
|
||||
shipping profile (docs/STATUS.md, reproducing the rate-control result). Costs
|
||||
no encode -- it reads the emitted bitstream, ~3 s.
|
||||
"""
|
||||
import sys, os, argparse
|
||||
sys.path.insert(0, "tools/encoder")
|
||||
import numpy as np
|
||||
from dlx import DLX
|
||||
|
||||
# FINDINGS 24: measured on the emulated 68000, instruction cycles only.
|
||||
BLIT_PCT, DIRECT_PCT = 53.6, 76.6
|
||||
CROSSOVER = 100 * BLIT_PCT / DIRECT_PCT
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("container", nargs="?",
|
||||
default="tmp/rc_fr_singe_sasi_rcprofile.dlx")
|
||||
ap.add_argument("--fix", default="none",
|
||||
choices=("none", "dual", "resync", "compose", "direct"))
|
||||
a = ap.parse_args()
|
||||
|
||||
if not os.path.exists(a.container):
|
||||
sys.exit(f"missing {a.container} -- see docs/STATUS.md, "
|
||||
f"'Reproducing the rate-control result'")
|
||||
|
||||
d = DLX(a.container)
|
||||
print(f"{a.container}: {d.W}x{d.H} {d.nframes} frames, {d.nb} blocks/frame, "
|
||||
f"k1={d.k1} k4={d.k4}")
|
||||
print(f"path choice: compose if non-SKIP > {CROSSOVER:.1f}% of blocks "
|
||||
f"(53.6% flat vs 76.6% x fraction), strategy={a.fix}\n")
|
||||
|
||||
# gv = what is on screen. ram = the player's RAM reference frame.
|
||||
# truth = what a correct player displays. All are palette-index canvases.
|
||||
gv = np.zeros((d.H, d.W), np.uint8)
|
||||
ram = np.zeros((d.H, d.W), np.uint8)
|
||||
truth = np.zeros((d.H, d.W), np.uint8)
|
||||
|
||||
def put(canvas, blks):
|
||||
for i, blk in blks.items():
|
||||
by, bx = divmod(i, d.nbx)
|
||||
canvas[by*4:by*4+4, bx*4:bx*4+4] = blk
|
||||
|
||||
drift_px, used, switches, resyncs = [], [], 0, 0
|
||||
prev_path = None
|
||||
for f in range(d.nframes):
|
||||
mode, blks = d.blocks(f)
|
||||
put(truth, blks)
|
||||
|
||||
frac = 100 * (mode != 0).mean()
|
||||
if a.fix == "compose": path = "compose"
|
||||
elif a.fix == "direct": path = "direct"
|
||||
else: path = "compose" if frac > CROSSOVER else "direct"
|
||||
|
||||
if path == "compose":
|
||||
if a.fix == "resync" and prev_path == "direct":
|
||||
ram = gv.copy() # re-read GVRAM into the RAM reference
|
||||
resyncs += 1
|
||||
put(ram, blks)
|
||||
gv = ram.copy() # full row-linear blit
|
||||
else:
|
||||
put(gv, blks)
|
||||
if a.fix == "dual":
|
||||
put(ram, blks) # keep the reference coherent as we go
|
||||
|
||||
if prev_path is not None and path != prev_path:
|
||||
switches += 1
|
||||
prev_path = path
|
||||
used.append(path)
|
||||
drift_px.append(int((gv != truth).sum()))
|
||||
|
||||
drift_px = np.array(drift_px)
|
||||
npx = d.H * d.W
|
||||
nc = used.count("compose")
|
||||
print(f"path used: compose {nc}/{d.nframes} ({100*nc/d.nframes:.0f}%), "
|
||||
f"direct {d.nframes-nc} -- {switches} switches between them"
|
||||
+ (f", {resyncs} resyncs" if resyncs else ""))
|
||||
bad = int((drift_px > 0).sum())
|
||||
print(f"\nframes displaying pixels no correct player would display: "
|
||||
f"{bad}/{d.nframes}")
|
||||
if bad:
|
||||
print(f" worst frame {drift_px.max()} px "
|
||||
f"({100*drift_px.max()/npx:.1f}% of the screen), "
|
||||
f"mean {drift_px.mean():.0f} px ({100*drift_px.mean()/npx:.1f}%)")
|
||||
first = int(np.argmax(drift_px > 0))
|
||||
print(f" first corrupt frame: {first} (path={used[first]}, "
|
||||
f"previous={used[first-1] if first else '-'})")
|
||||
|
||||
# Cost of the strategy, in % of a 12fps frame budget. 'dual' pays 1.52x on the
|
||||
# direct path: the same block written twice, +108 cycles on 208 (FINDINGS 28.2).
|
||||
mult = 1.52 if a.fix == "dual" else 1.0
|
||||
cost = np.array([BLIT_PCT if p == "compose" else DIRECT_PCT * mult *
|
||||
(d.modes(f) != 0).mean()
|
||||
for f, p in enumerate(used)])
|
||||
if a.fix == "resync":
|
||||
for f in range(1, d.nframes):
|
||||
if used[f] == "compose" and used[f-1] == "direct":
|
||||
cost[f] += BLIT_PCT # the GVRAM->RAM re-read is a full frame
|
||||
print(f"\ndisplay cost: median {np.median(cost):.1f}% "
|
||||
f"p90 {np.percentile(cost,90):.1f}% max {cost.max():.1f}% "
|
||||
f"of a 12fps frame budget")
|
||||
over = int((cost > 100).sum())
|
||||
if over:
|
||||
print(f" frames that do NOT fit in the budget at all: {over}/{d.nframes}")
|
||||
|
||||
sys.exit(1 if bad else 0)
|
||||
@@ -0,0 +1,104 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Per-frame CPU cost of the real decoder, from MEASURED per-mode block costs.
|
||||
|
||||
python3 tools/analysis/11_cpu_budget.py [container.dlx]
|
||||
|
||||
FINDINGS 24.5 priced the display path as "76.6% of a 12fps frame x the non-SKIP
|
||||
block fraction", i.e. every non-SKIP block costs the same. It does not: the four
|
||||
block modes were measured separately on the 68000 (synthetic single-mode frames,
|
||||
tools/bench/prep_dlx.py) and V4 costs 1.5x V1. Since V4 is roughly half of all
|
||||
non-SKIP blocks on hard content, the old model runs ~1.8x optimistic exactly
|
||||
where it matters.
|
||||
|
||||
This applies the measured costs to a real container's mode histograms and
|
||||
reports what fraction of frames actually fit 833,333 cycles.
|
||||
|
||||
Costs are MEASURED (tools/bench/decode.lua), cross-checked against hand-derived
|
||||
MC68000 timings in FINDINGS 28.4. They are instruction cycles against
|
||||
zero-wait-state memory, so like every figure in this project since FINDINGS 24
|
||||
they are a LOWER BOUND -- real GVRAM stalls the CPU.
|
||||
"""
|
||||
import sys, os, argparse
|
||||
sys.path.insert(0, "tools/encoder")
|
||||
import numpy as np
|
||||
from dlx import DLX
|
||||
|
||||
CPUHZ = 10_000_000
|
||||
FPS = 12
|
||||
FRAME = CPUHZ / FPS # 833,333 cycles
|
||||
|
||||
# cycles per block, measured on the emulated 68000 (synthetic single-mode frames)
|
||||
C_V1, C_V4, C_RAW = 299.9, 448.2, 400.4
|
||||
C_SKIP_FAST = 53.0 / 4 # all-SKIP header byte: one tst.b for 4
|
||||
C_SKIP_MIXED = 45.0 # a SKIP block inside a mixed byte
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("container", nargs="?",
|
||||
default="tmp/rc_fr_singe_sasi_rcprofile.dlx")
|
||||
a = ap.parse_args()
|
||||
if not os.path.exists(a.container):
|
||||
sys.exit(f"missing {a.container}")
|
||||
|
||||
d = DLX(a.container)
|
||||
|
||||
def cycles(mode):
|
||||
g = mode.reshape(-1, 4) # one header byte = four blocks
|
||||
allskip = (g == 0).all(1)
|
||||
c = allskip.sum() * 4 * C_SKIP_FAST
|
||||
m = g[~allskip]
|
||||
c += (m == 0).sum() * C_SKIP_MIXED
|
||||
c += (m == 1).sum() * C_V1
|
||||
c += (m == 2).sum() * C_V4
|
||||
c += (m == 3).sum() * C_RAW
|
||||
return c
|
||||
|
||||
modes = [d.modes(f) for f in range(d.nframes)]
|
||||
cyc = np.array([cycles(m) for m in modes])
|
||||
pct = 100 * cyc / FRAME
|
||||
ns = np.array([100 * (m != 0).mean() for m in modes])
|
||||
|
||||
print(f"{a.container}: {d.nframes} frames, {d.nb} blocks/frame")
|
||||
print(f"measured block costs: SKIP {C_SKIP_FAST*4:.0f}/4 (clustered) "
|
||||
f"{C_SKIP_MIXED:.0f} (mixed) V1 {C_V1:.0f} V4 {C_V4:.0f} RAW {C_RAW:.0f} cycles\n")
|
||||
|
||||
# --- validation against the four real frames timed on the 68000. These
|
||||
# timings belong to ONE container; quoting them against any other would be
|
||||
# comparing a model of this stream to a measurement of a different one.
|
||||
TIMED = "tmp/rc_fr_singe_sasi_rcprofile.dlx"
|
||||
TIMED_FRAMES = (("min non-SKIP", 15.4, 31.5), ("median", 48.1, 73.8),
|
||||
("p90", 82.5, 116.4), ("max non-SKIP", 100.0, 135.8))
|
||||
if os.path.abspath(a.container) == os.path.abspath(TIMED):
|
||||
print("model vs the frames actually timed on the 68000:")
|
||||
for label, frac, meas in TIMED_FRAMES:
|
||||
i = int(np.argmin(abs(ns - frac)))
|
||||
print(f" {label:<14} non-SKIP {ns[i]:5.1f}% model {pct[i]:6.1f}% "
|
||||
f"measured {meas:5.1f}% error {pct[i]-meas:+.1f} pt")
|
||||
else:
|
||||
print(f"(no 68000 timings for this container -- the model was validated to "
|
||||
f"within\n 1 pt on {TIMED}; run tools/bench/decode.lua to time this one)")
|
||||
|
||||
old = 76.6 * ns / 100
|
||||
print(f"\nper-frame cost, % of a {FPS}fps frame budget:")
|
||||
print(f" measured-cost model: median {np.median(pct):5.1f} "
|
||||
f"p90 {np.percentile(pct,90):5.1f} max {pct.max():5.1f}")
|
||||
print(f" FINDINGS 24.5 model: median {np.median(old):5.1f} "
|
||||
f"p90 {np.percentile(old,90):5.1f} max {old.max():5.1f} "
|
||||
f"(optimistic by {np.median(pct)/np.median(old):.2f}x at the median)")
|
||||
|
||||
miss = pct > 100
|
||||
print(f"\nframes that do NOT fit 833,333 cycles: {miss.sum()}/{d.nframes} "
|
||||
f"({100*miss.mean():.0f}%)")
|
||||
if miss.any():
|
||||
print(f" worst {pct.max():.1f}% -- {(pct.max()-100)/100*1000/FPS:.0f} ms late "
|
||||
f"on an {1000/FPS:.0f} ms frame")
|
||||
print(f" the budget is first missed at {ns[miss].min():.1f}% non-SKIP blocks")
|
||||
|
||||
# Where do the cycles go? This is what a cost-aware mode decision would act on.
|
||||
tot = np.array([[(m == k).sum() for k in range(4)] for m in modes]).sum(0)
|
||||
spend = tot * np.array([C_SKIP_MIXED, C_V1, C_V4, C_RAW])
|
||||
print(f"\nwhere the cycles go, over the whole window:")
|
||||
for k, n in enumerate(("SKIP", "V1", "V4", "RAW")):
|
||||
print(f" {n:<5} {100*tot[k]/tot.sum():5.1f}% of blocks "
|
||||
f"{100*spend[k]/spend.sum():5.1f}% of the cycles")
|
||||
print(f"\nV4 is {C_V4/C_V1:.2f}x a V1 block for {4}x the payload bytes -- the mode "
|
||||
f"decision\nin vq_hybrid.py charges it the bytes but not the cycles.")
|
||||
@@ -40,4 +40,34 @@ python3 tools/analysis/09_ratectl_drift.py > tmp/drift_check.log 2>&1 \
|
||||
|| { cat tmp/drift_check.log; exit 1; }
|
||||
tail -9 tmp/drift_check.log
|
||||
|
||||
echo "--- session 7: display-path coherency (FINDINGS 28.1) ---"
|
||||
# 10_pathmix_drift.py is a COUNTEREXAMPLE, kept runnable: the dual-path plan of
|
||||
# FINDINGS 24.5/25.6 must still be shown to corrupt frames, and the strategy the
|
||||
# player actually uses must still be clean. A green light here means the reason
|
||||
# decode.s has one display path is still demonstrable, not just asserted.
|
||||
python3 tools/analysis/10_pathmix_drift.py > tmp/pathmix.log 2>&1 \
|
||||
&& { echo "FAIL: the dual-path plan no longer reproduces its own defect"; \
|
||||
cat tmp/pathmix.log; exit 1; }
|
||||
grep -a "frames displaying pixels" tmp/pathmix.log
|
||||
python3 tools/analysis/10_pathmix_drift.py --fix direct > tmp/pathmix_direct.log 2>&1 \
|
||||
|| { echo "FAIL: direct-to-GVRAM is no longer coherent"; cat tmp/pathmix_direct.log; exit 1; }
|
||||
|
||||
echo "--- session 7: 68000 decoder is pixel-exact (FINDINGS 28) ---"
|
||||
# The strongest display test in the tree: 120 frames decoded in sequence by
|
||||
# 68000 code, every block mode, full temporal recursion. A SKIP block is a claim
|
||||
# about the previous frame still being on screen, so the last frame is only
|
||||
# right if all 120 were.
|
||||
DLX=tmp/rc_fr_singe_sasi_rcprofile.dlx
|
||||
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile sasi
|
||||
python3 tools/bench/prep_dlx.py "$DLX" > tmp/prep_dlx.log
|
||||
tools/vasm/vasmm68k_mot -Fbin -o tmp/decode.bin src/player/decode.s > /dev/null
|
||||
mkdir -p tmp/snap_decode
|
||||
rm -f tmp/snap_decode/x68000/*.png
|
||||
( cd tmp && DLX_VERIFY_ONLY=1 SDL_VIDEODRIVER=dummy timeout -k 5 300 mame x68000 \
|
||||
-bios ipl10 -ramsize 2M -video soft -window -sound none -nothrottle -plugins \
|
||||
-autoboot_script ../tools/bench/decode.lua \
|
||||
-snapshot_directory ./snap_decode -snapview native -seconds_to_run 20 \
|
||||
> decode_check.log 2>&1 )
|
||||
python3 tools/bench/verify_decode.py "$DLX"
|
||||
|
||||
echo "ALL GREEN"
|
||||
|
||||
@@ -0,0 +1,172 @@
|
||||
-- Time and verify src/player/decode.s on the emulated 68000.
|
||||
--
|
||||
-- Two questions, one run:
|
||||
-- 1. CORRECTNESS. Decode the whole window frame by frame and snapshot the
|
||||
-- last frame. tools/bench/verify_decode.py checks it against the Python
|
||||
-- reference decoder (tools/encoder/dlx.py) pixel-for-pixel. Every SKIP
|
||||
-- block in every frame is a claim about the previous frame still being on
|
||||
-- screen, so a sequential run is the only honest test -- decoding one
|
||||
-- frame in isolation would prove nothing about the temporal recursion.
|
||||
-- 2. COST. Time individual frames chosen across the non-SKIP distribution,
|
||||
-- not its mean (FINDINGS 25.6), plus one full 120-frame pass.
|
||||
--
|
||||
-- MEASUREMENT SCOPE, unchanged from blit.lua: MAME's gvram_w/gvram_r carry no
|
||||
-- timing at all, so these are pure 68000 instruction cycles against
|
||||
-- zero-wait-state memory -- a LOWER BOUND on real hardware, not a prediction.
|
||||
-- Interrupts are masked (SR=$2700) so the IPL cannot steal cycles.
|
||||
--
|
||||
-- Codebook expansion and palette packing are done host-side by prep_dlx.py:
|
||||
-- they are load-time costs, not per-frame ones, and including them would
|
||||
-- flatter or damn the inner loop for no reason.
|
||||
|
||||
M = manager.machine
|
||||
SP = M.devices[":maincpu"].spaces["program"]
|
||||
|
||||
local function findfile(n)
|
||||
for _,p in ipairs{"../tools/bench/"..n, "tools/bench/"..n, n} do
|
||||
local f = io.open(p,"rb"); if f then f:close(); return p end
|
||||
end
|
||||
error(n.." not found")
|
||||
end
|
||||
local MODE = loadfile(findfile("crtc_mode.lua"))()
|
||||
local META = loadfile("decode_meta.lua")()
|
||||
|
||||
local FLAG, ITER, NFR, FPTR = 0x18000, 0x18008, 0x1800C, 0x18010
|
||||
local CB1, CB4, STREAM = 0x20000, 0x22000, 0x30000
|
||||
local GVRAM, GPAL = 0xC00000, 0xE82000
|
||||
local CPUHZ = 10000000 -- x68k.cpp:1133, 40_MHz_XTAL/4
|
||||
local FRAME12 = CPUHZ / META.fps
|
||||
|
||||
local code do local f=io.open("decode.bin","rb"); code=f:read("a"); f:close() end
|
||||
local data do local f=io.open("decode_data.bin","rb"); data=f:read("a"); f:close() end
|
||||
|
||||
local YOFF = (MODE.height - META.H) // 2
|
||||
local function T() local t=M.time; return t.seconds + t.attoseconds/1e18 end
|
||||
local function P(s) print("[DEC] "..s) end
|
||||
|
||||
-- Bulk-load a slice of the blob as big-endian longwords. 1 MB one byte at a
|
||||
-- time is 1M Lua->C calls; longwords cut that by four.
|
||||
local function push(addr, s, from, len)
|
||||
local i, n = from, len
|
||||
while n >= 4 do
|
||||
SP:write_u32(addr, (string.unpack(">I4", s, i)))
|
||||
addr, i, n = addr+4, i+4, n-4
|
||||
end
|
||||
while n > 0 do
|
||||
SP:write_u8(addr, string.byte(s,i)); addr, i, n = addr+1, i+1, n-1
|
||||
end
|
||||
end
|
||||
|
||||
local function setup()
|
||||
MODE.apply(SP)
|
||||
local o = 1
|
||||
push(CB1, data, o, META.cb1_len); o = o + META.cb1_len
|
||||
push(CB4, data, o, META.cb4_len); o = o + META.cb4_len
|
||||
local palo = o; o = o + META.pal_len
|
||||
push(STREAM, data, o, META.stream_len)
|
||||
for c = 0, 255 do
|
||||
SP:write_u16(GPAL + c*2, (string.unpack(">I2", data, palo + c*2)))
|
||||
end
|
||||
-- Active area starts at index 0, exactly as the reference decoder's canvas
|
||||
-- does; the letterbox gets the palette's darkest entry because the encoder
|
||||
-- does not yet reserve a black one (docs/STATUS.md, encoder gaps).
|
||||
for y = 0, MODE.height-1 do
|
||||
local base, v = GVRAM + y*1024, 0
|
||||
if y < YOFF or y >= YOFF+META.H then v = META.dark end
|
||||
for x = 0, MODE.width-1 do SP:write_u16(base + x*2, v) end
|
||||
end
|
||||
for i = 1, #code do SP:write_u8(0x10000+i-1, string.byte(code,i)) end
|
||||
P(string.format("loaded decode.bin=%d B, codebooks %d+%d B, stream %d B, %d frames",
|
||||
#code, META.cb1_len, META.cb4_len, META.stream_len, META.nframes))
|
||||
end
|
||||
|
||||
local function launch(off, nfr, iter)
|
||||
SP:write_u32(FLAG, 0)
|
||||
SP:write_u32(ITER, iter)
|
||||
SP:write_u32(NFR, nfr)
|
||||
SP:write_u32(FPTR, STREAM + off)
|
||||
local cpu = M.devices[":maincpu"]
|
||||
cpu.state["SR"].value = 0x2700 -- supervisor, ALL interrupts masked
|
||||
cpu.state["SP"].value = 0x8000
|
||||
cpu.state["PC"].value = 0x10000
|
||||
end
|
||||
|
||||
-- The plan: one sequential correctness pass, then the cost anchors, then a
|
||||
-- full pass timed. Iteration counts target ~4 emulated seconds each so the
|
||||
-- 1/55.46 s timing granularity costs under 0.5%.
|
||||
-- DLX_VERIFY_ONLY=1 drops the cost anchors and runs only the correctness pass,
|
||||
-- so tools/bench/check.sh can gate the decoder without paying for ~2 minutes of
|
||||
-- timing runs that would make the green light sensitive to host load anyway.
|
||||
local VERIFY_ONLY = os.getenv("DLX_VERIFY_ONLY") == "1"
|
||||
|
||||
local PLAN = { {name="sequential decode of all "..META.nframes.." frames (correctness)",
|
||||
off=0, nfr=META.nframes, iter=1, snap=true} }
|
||||
for _,an in ipairs(VERIFY_ONLY and {} or META.anchors) do
|
||||
local est = math.max(0.06, an.frac/100) * 1.30 * FRAME12
|
||||
PLAN[#PLAN+1] = {name="frame @ "..an.name, off=an.off, nfr=1,
|
||||
iter=math.max(20, math.floor(4*CPUHZ/est)), frac=an.frac}
|
||||
end
|
||||
if not VERIFY_ONLY then
|
||||
PLAN[#PLAN+1] = {name="full "..META.nframes.."-frame pass (mean over the window)",
|
||||
off=0, nfr=META.nframes, iter=1, seq=true}
|
||||
end
|
||||
|
||||
local step, st, t0 = 0, "boot", nil
|
||||
local results = {}
|
||||
|
||||
local function report(p, dt)
|
||||
local per = p.nfr * p.iter
|
||||
local cyc = dt * CPUHZ / per
|
||||
local pct = 100 * cyc / FRAME12
|
||||
if p.snap then return end -- correctness pass, iter=1, too coarse
|
||||
results[#results+1] = {p=p, cyc=cyc, pct=pct}
|
||||
P(string.format("%s", p.name))
|
||||
P(string.format(" %d frames in %.4f s -> %.0f cycles/frame = %.1f%% of a %dfps frame",
|
||||
per, dt, cyc, pct, META.fps))
|
||||
end
|
||||
|
||||
SUB = emu.add_machine_frame_notifier(function()
|
||||
local ok, err = pcall(function()
|
||||
local t = T()
|
||||
if st == "boot" then
|
||||
if t < 3.0 then return end
|
||||
setup(); step = 1; launch(PLAN[1].off, PLAN[1].nfr, PLAN[1].iter)
|
||||
st, t0 = "running", nil; return
|
||||
end
|
||||
if st == "running" then
|
||||
local fl = SP:read_u32(FLAG)
|
||||
if fl == 1 and not t0 then t0 = t; return end
|
||||
if fl == 0xEE then
|
||||
P("BITSTREAM DESYNC -- decoder consumed the wrong number of payload bytes")
|
||||
M:exit(); return
|
||||
end
|
||||
if fl == 0xFF then
|
||||
report(PLAN[step], t - (t0 or t))
|
||||
if PLAN[step].snap then st = "snap"; return end
|
||||
step = step + 1
|
||||
if PLAN[step] then
|
||||
launch(PLAN[step].off, PLAN[step].nfr, PLAN[step].iter)
|
||||
st, t0 = "running", nil
|
||||
else st = "finish" end
|
||||
return
|
||||
end
|
||||
if t > 400 then P("TIMEOUT flag="..string.format("%08X",fl)); M:exit() end
|
||||
return
|
||||
end
|
||||
if st == "snap" then
|
||||
M.video:snapshot()
|
||||
P("snapshot taken after the sequential pass -- last frame, 68000-decoded")
|
||||
step = step + 1
|
||||
launch(PLAN[step].off, PLAN[step].nfr, PLAN[step].iter)
|
||||
st, t0 = "running", nil; return
|
||||
end
|
||||
if st == "finish" then
|
||||
P("---- summary (instruction cycles only; real GVRAM adds wait states) ----")
|
||||
for _,r in ipairs(results) do
|
||||
P(string.format(" %-46s %8.0f cyc %5.1f%% of a frame", r.p.name, r.cyc, r.pct))
|
||||
end
|
||||
M:exit()
|
||||
end
|
||||
end)
|
||||
if not ok then print("[DEC] LUA ERROR: "..tostring(err)); M:exit() end
|
||||
end)
|
||||
@@ -0,0 +1,123 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Lay a DLX1 container out the way src/player/decode.s expects to find it.
|
||||
|
||||
python3 tools/bench/prep_dlx.py <in.dlx> [--out tmp/decode]
|
||||
|
||||
Writes <out>_data.bin (one blob Lua pushes into emulated RAM) and <out>_meta.lua
|
||||
(sizes, per-frame record offsets, and the timing anchors).
|
||||
|
||||
Two things happen here that the shipping player would do at load time on the
|
||||
68000 itself, and are therefore NOT part of the per-frame cost being measured:
|
||||
|
||||
* codebook expansion to word-per-pixel form. CB1 -> 32 B/entry, CB4 -> 8 B,
|
||||
so the inner loop scales an index with lsl.w #5 / #3 and movems the result
|
||||
straight into GVRAM with no unpacking. 8 KB + 2 KB of the 2 MB.
|
||||
* palette packing to GGGGGRRRRRBBBBBI with the shared LSB I chosen PER ENTRY
|
||||
by minimum squared error (FINDINGS 23.3, worth 1.96 dB).
|
||||
|
||||
The encoder still emits 24-bit palettes and does not reserve a black entry
|
||||
(known gap, docs/STATUS.md), so the letterbox here is filled with whatever
|
||||
palette entry is closest to black rather than a true reserved black. That is
|
||||
cosmetic and outside the active 256x192 area the decoder is judged on.
|
||||
|
||||
A synthetic all-SKIP frame is appended to the stream. No real frame is all
|
||||
SKIP, but it prices the mode-header walk on its own -- the per-block cost the
|
||||
"76.6% x non-SKIP fraction" model in FINDINGS 24.5 leaves out entirely.
|
||||
"""
|
||||
import sys, os, argparse
|
||||
sys.path.insert(0, "tools/encoder")
|
||||
import numpy as np
|
||||
from dlx import DLX
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("container")
|
||||
ap.add_argument("--out", default="tmp/decode")
|
||||
a = ap.parse_args()
|
||||
|
||||
d = DLX(a.container)
|
||||
if d.idx_bytes != 1:
|
||||
sys.exit("2-byte codebook indices: decode.s assumes 1 (k<=256)")
|
||||
|
||||
# --- codebooks, expanded to one WORD per pixel (high byte is discarded by
|
||||
# gvram_w, so it is left zero and never has to be cleared)
|
||||
cb1 = np.zeros((d.k1, 16, 2), np.uint8); cb1[:, :, 1] = d.cb1.reshape(d.k1, 16)
|
||||
cb4 = np.zeros((d.k4, 4, 2), np.uint8); cb4[:, :, 1] = d.cb4.reshape(d.k4, 4)
|
||||
|
||||
# --- palette words, I chosen per entry (identical maths to verify_frame256.py)
|
||||
pal = d.pal.astype(int)
|
||||
p6 = lambda v: ((v << 2) | (v >> 4)) & 0xFF
|
||||
f = pal >> 3
|
||||
render = lambda I: p6((f << 1) | I[:, None])
|
||||
I = (((render(np.ones(256, int)) - pal) ** 2).sum(1)
|
||||
< ((render(np.zeros(256, int)) - pal) ** 2).sum(1)).astype(int)
|
||||
words = (f[:, 1] << 11) | (f[:, 0] << 6) | (f[:, 2] << 1) | I
|
||||
palb = np.zeros((256, 2), np.uint8)
|
||||
palb[:, 0], palb[:, 1] = words >> 8, words & 0xFF
|
||||
dark = int(((render(I).astype(int)) ** 2).sum(1).argmin())
|
||||
|
||||
# --- frame stream: [u32 len][modes][payload] per frame, each record start
|
||||
# rounded up to a 4-byte boundary.
|
||||
#
|
||||
# This padding is not cosmetic. Payload lengths are arbitrary, so laid end
|
||||
# to end the records land on odd addresses, and `move.l (a0)+` at an odd
|
||||
# address is an ADDRESS ERROR on a 68000 -- it vectors into the IPL rather
|
||||
# than reading slowly. The container as written by encode.py is unaligned,
|
||||
# so this loader realigns it; the encoder should carry the padding itself
|
||||
# (FINDINGS 28.3). It costs at most 3 bytes per frame -- 36 B/s at 12fps,
|
||||
# against a 110 KB/s budget.
|
||||
stream, rec_off, pad = bytearray(), [], 0
|
||||
for (o, n) in d.frames:
|
||||
while len(stream) % 4:
|
||||
stream += b"\0"; pad += 1
|
||||
rec_off.append(len(stream))
|
||||
stream += n.to_bytes(4, "big") + d.raw[o:o + n]
|
||||
|
||||
# Synthetic single-mode frames. No real frame is all one mode, but the mix is
|
||||
# exactly what the "76.6% x non-SKIP fraction" model of FINDINGS 24.5 assumes
|
||||
# away: it prices every non-SKIP block as one V1-style burst. These four price
|
||||
# the modes separately, which is the only way to see which one is expensive.
|
||||
synth = {}
|
||||
for name, mo, per in (("all-SKIP", 0, 0), ("all-V1", 1, 1),
|
||||
("all-V4", 2, 4), ("all-RAW", 3, 16)):
|
||||
while len(stream) % 4:
|
||||
stream += b"\0"; pad += 1
|
||||
synth[name] = len(stream)
|
||||
hdr = bytes([mo * 0x55] * d.mode_bytes)
|
||||
stream += (d.mode_bytes + d.nb * per).to_bytes(4, "big") + hdr + bytes(d.nb * per)
|
||||
|
||||
# --- timing anchors: the distribution, not its mean (FINDINGS 25.6's lesson)
|
||||
ns = np.array([100 * (d.modes(i) != 0).mean() for i in range(d.nframes)])
|
||||
order = np.argsort(ns)
|
||||
pick = {
|
||||
"min non-SKIP %.1f%%" % ns[order[0]]: int(order[0]),
|
||||
"median %.1f%%" % np.median(ns): int(order[len(order)//2]),
|
||||
"p90 %.1f%%" % ns[order[int(.9*len(order))]]: int(order[int(.9*len(order))]),
|
||||
"max non-SKIP %.1f%%" % ns[order[-1]]: int(order[-1]),
|
||||
}
|
||||
anchors = [(n, rec_off[i], float(ns[i])) for n, i in pick.items()]
|
||||
for name in ("all-SKIP", "all-V1", "all-V4", "all-RAW"):
|
||||
anchors.append((f"synthetic {name}", synth[name],
|
||||
0.0 if name == "all-SKIP" else 100.0))
|
||||
|
||||
blob = cb1.tobytes() + cb4.tobytes() + palb.tobytes() + bytes(stream)
|
||||
open(a.out + "_data.bin", "wb").write(blob)
|
||||
|
||||
with open(a.out + "_meta.lua", "w") as fh:
|
||||
fh.write("-- generated by tools/bench/prep_dlx.py -- do not edit\nreturn {\n")
|
||||
fh.write(f" W={d.W}, H={d.H}, fps={d.fps}, nframes={d.nframes},\n")
|
||||
fh.write(f" k1={d.k1}, k4={d.k4}, dark={dark},\n")
|
||||
fh.write(f" cb1_len={cb1.nbytes}, cb4_len={cb4.nbytes}, pal_len={palb.nbytes},\n")
|
||||
fh.write(f" stream_len={len(stream)},\n")
|
||||
fh.write(" anchors={\n")
|
||||
for n, o, frac in anchors:
|
||||
fh.write(f' {{name="{n}", off={o}, frac={frac:.1f}}},\n')
|
||||
fh.write(" },\n}\n")
|
||||
|
||||
print(f"{a.container}: {d.nframes} frames, {d.W}x{d.H}, k1={d.k1} k4={d.k4}")
|
||||
print(f" cb1 {cb1.nbytes} B + cb4 {cb4.nbytes} B expanded, palette {palb.nbytes} B, "
|
||||
f"stream {len(stream)} B -> {a.out}_data.bin ({len(blob)} B)")
|
||||
print(f" non-SKIP blocks/frame: median {np.median(ns):.1f}% "
|
||||
f"p90 {np.percentile(ns,90):.1f}% max {ns.max():.1f}%")
|
||||
print(f" darkest palette entry: index {dark} -> {tuple(render(I)[dark])}")
|
||||
print(f" 4-byte record alignment cost {pad} B over {d.nframes} frames "
|
||||
f"({pad / d.nframes:.2f} B/frame = {pad / d.nframes * d.fps:.0f} B/s)")
|
||||
@@ -0,0 +1,66 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Is the 68000 decoder's output pixel-exact against the reference decoder?
|
||||
|
||||
python3 tools/bench/verify_decode.py <in.dlx> [--snap tmp/snap_decode]
|
||||
|
||||
Checks tmp/snap_decode/x68000/0000.png -- the screen after src/player/decode.s
|
||||
has decoded every frame of the container in sequence -- against
|
||||
tools/encoder/dlx.py's reconstruction of the final frame.
|
||||
|
||||
This is a stronger test than the blit regression it is modelled on. The blit
|
||||
proved the 68000 could COPY a frame; this proves it can PARSE one. And because
|
||||
the decoder is temporally recursive -- a SKIP block is a claim that the previous
|
||||
frame is still in GVRAM -- the last frame of a sequential run is only correct if
|
||||
every frame before it was, so a single comparison audits all of them.
|
||||
"""
|
||||
import argparse, sys
|
||||
sys.path.insert(0, "tools/encoder")
|
||||
import numpy as np
|
||||
from PIL import Image
|
||||
from dlx import DLX
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("container")
|
||||
ap.add_argument("--snap", default="tmp/snap_decode")
|
||||
a = ap.parse_args()
|
||||
|
||||
d = DLX(a.container)
|
||||
canvas = np.zeros((d.H, d.W), np.uint8)
|
||||
for f in range(d.nframes):
|
||||
d.paint(canvas, f)
|
||||
|
||||
pal = d.pal.astype(int)
|
||||
p6 = lambda v: ((v << 2) | (v >> 4)) & 0xFF
|
||||
fl = pal >> 3
|
||||
render = lambda I: p6((fl << 1) | I[:, None])
|
||||
I = (((render(np.ones(256, int)) - pal) ** 2).sum(1)
|
||||
< ((render(np.zeros(256, int)) - pal) ** 2).sum(1)).astype(int)
|
||||
exp = render(I)[canvas]
|
||||
|
||||
s = np.asarray(Image.open(f"{a.snap}/x68000/0000.png").convert("RGB")).astype(int)
|
||||
fail = []
|
||||
if s.shape[:2] != (512, 256):
|
||||
fail.append(f"1. geometry: expected 512x256, got {s.shape[1]}x{s.shape[0]}")
|
||||
else:
|
||||
if not all(np.array_equal(s[i], s[i+1]) for i in range(1, s.shape[0]-1, 2)):
|
||||
fail.append("2. double-scan pairing (1,2),(3,4),... broken")
|
||||
g = s[0::2]
|
||||
yoff = (g.shape[0] - d.H) // 2
|
||||
act = g[yoff:yoff+d.H]
|
||||
if not np.array_equal(act, exp):
|
||||
diff = abs(act - exp)
|
||||
bad = diff.any(2)
|
||||
by, bx = np.where(bad)
|
||||
blocks = sorted(set(zip((by//4).tolist(), (bx//4).tolist())))
|
||||
fail.append(f"3. frame {d.nframes-1} not pixel-exact: {bad.sum()} px in "
|
||||
f"{len(blocks)} blocks differ, maxdiff {diff.max()}; "
|
||||
f"first block (by={blocks[0][0]}, bx={blocks[0][1]})")
|
||||
|
||||
for x in fail:
|
||||
print("FAIL " + x)
|
||||
if fail:
|
||||
sys.exit(1)
|
||||
print(f"OK {d.nframes} frames decoded on the 68000, final frame pixel-exact "
|
||||
f"against tools/encoder/dlx.py")
|
||||
print(f" {d.W}x{d.H}, {d.nb} blocks/frame, k1={d.k1} k4={d.k4}, "
|
||||
f"all four block modes exercised")
|
||||
@@ -0,0 +1,112 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Reference DLX1 reader/decoder -- the ground truth the 68000 player is checked against.
|
||||
|
||||
This is deliberately a *decoder*, not a re-run of the encoder: it parses the
|
||||
container byte-for-byte the way `src/player/` must, so that any disagreement
|
||||
between it and the 68000 is a decoder bug rather than two encoders differing.
|
||||
`tools/analysis/09_ratectl_drift.py` already gates the encoder against its own
|
||||
reconstruction; this gates the container against the player.
|
||||
|
||||
Everything is big-endian (see the `encode.py` docstring). Block raster order,
|
||||
2-bit modes packed MSB-first: 00=SKIP 01=V1 10=V4 11=RAW.
|
||||
|
||||
V4 sub-block order is (sub_y, sub_x) row-major -- TL, TR, BL, BR -- matching
|
||||
`vq_hybrid.paint`'s reshape(-1,2,2,2,2).transpose(0,1,3,2,4).
|
||||
"""
|
||||
import struct
|
||||
import numpy as np
|
||||
|
||||
MODE_SKIP, MODE_V1, MODE_V4, MODE_RAW = 0, 1, 2, 3
|
||||
|
||||
|
||||
class DLX:
|
||||
def __init__(self, path):
|
||||
self.raw = open(path, "rb").read()
|
||||
b = self.raw
|
||||
if b[:4] != b"DLX1":
|
||||
raise ValueError(f"{path}: not a DLX1 container")
|
||||
(self.W, self.H, self.fps, self.nframes,
|
||||
self.k1, self.k4) = struct.unpack(">HHHHHH", b[4:16])
|
||||
off_pal, off_cb1, off_cb4, off_frm = struct.unpack(">IIII", b[16:32])
|
||||
|
||||
self.pal = np.frombuffer(b, np.uint8, 256 * 3, off_pal).reshape(256, 3)
|
||||
self.cb1 = np.frombuffer(b, np.uint8, self.k1 * 16,
|
||||
off_cb1).reshape(self.k1, 4, 4)
|
||||
self.cb4 = np.frombuffer(b, np.uint8, self.k4 * 4,
|
||||
off_cb4).reshape(self.k4, 2, 2)
|
||||
|
||||
self.idx_bytes = 1 if max(self.k1, self.k4) <= 256 else 2
|
||||
self.nbx, self.nby = self.W // 4, self.H // 4
|
||||
self.nb = self.nbx * self.nby
|
||||
self.mode_bytes = (self.nb * 2 + 7) // 8
|
||||
|
||||
# frame directory: (offset of the mode header, payload length)
|
||||
self.frames = []
|
||||
p = off_frm
|
||||
for _ in range(self.nframes):
|
||||
(n,) = struct.unpack(">I", b[p:p + 4])
|
||||
self.frames.append((p + 4, n))
|
||||
p += 4 + n
|
||||
if p != len(b):
|
||||
raise ValueError(f"{path}: {len(b) - p} trailing bytes after "
|
||||
f"{self.nframes} frames")
|
||||
|
||||
def modes(self, f):
|
||||
o, _ = self.frames[f]
|
||||
h = np.frombuffer(self.raw, np.uint8, self.mode_bytes, o)
|
||||
m = np.stack([(h >> 6) & 3, (h >> 4) & 3, (h >> 2) & 3, h & 3], axis=1)
|
||||
return m.reshape(-1)[:self.nb].copy()
|
||||
|
||||
def blocks(self, f):
|
||||
"""Decoded 4x4 palette-index blocks for the non-SKIP blocks of frame f.
|
||||
|
||||
Returns (mode, dict{block index -> (4,4) uint8}). SKIP blocks are
|
||||
absent by construction -- the player must leave those pixels alone,
|
||||
and a decoder that materialises them is hiding the very bug this
|
||||
module exists to catch.
|
||||
"""
|
||||
mode = self.modes(f)
|
||||
o, n = self.frames[f]
|
||||
p, end = o + self.mode_bytes, o + n
|
||||
ib, out = self.idx_bytes, {}
|
||||
b = self.raw
|
||||
for i, mo in enumerate(mode):
|
||||
if mo == MODE_SKIP:
|
||||
continue
|
||||
if mo == MODE_V1:
|
||||
v = b[p] if ib == 1 else (b[p] << 8) | b[p + 1]
|
||||
p += ib
|
||||
out[i] = self.cb1[v]
|
||||
elif mo == MODE_V4:
|
||||
sub = []
|
||||
for _ in range(4):
|
||||
v = b[p] if ib == 1 else (b[p] << 8) | b[p + 1]
|
||||
p += ib
|
||||
sub.append(self.cb4[v])
|
||||
blk = np.empty((4, 4), np.uint8)
|
||||
blk[0:2, 0:2], blk[0:2, 2:4] = sub[0], sub[1]
|
||||
blk[2:4, 0:2], blk[2:4, 2:4] = sub[2], sub[3]
|
||||
out[i] = blk
|
||||
else:
|
||||
out[i] = np.frombuffer(b, np.uint8, 16, p).reshape(4, 4)
|
||||
p += 16
|
||||
if p != end:
|
||||
raise ValueError(f"frame {f}: payload consumed {p - o} of {n} bytes")
|
||||
return mode, out
|
||||
|
||||
def paint(self, canvas, f):
|
||||
"""Apply frame f in place to a (H,W) index canvas. SKIP = untouched."""
|
||||
mode, blks = self.blocks(f)
|
||||
for i, blk in blks.items():
|
||||
by, bx = divmod(i, self.nbx)
|
||||
canvas[by * 4:by * 4 + 4, bx * 4:bx * 4 + 4] = blk
|
||||
return mode
|
||||
|
||||
def decode_all(self):
|
||||
"""The true reconstruction sequence: what any correct player displays."""
|
||||
c = np.zeros((self.H, self.W), np.uint8)
|
||||
out = []
|
||||
for f in range(self.nframes):
|
||||
self.paint(c, f)
|
||||
out.append(c.copy())
|
||||
return out
|
||||
Reference in New Issue
Block a user