The 68000 decoder draws pixel-exact frames, and does not fit
src/player/decode.s parses DLX1 and decodes straight into GVRAM. Verified pixel-exact over a 120-frame sequential run of the worst sustained window on the disc -- all four block modes, full temporal recursion, so the last frame is only right if all 120 were. In check.sh. It costs a mean of 81.7% of a 12fps frame budget, and 31% of frames exceed 100% (42% at scsi). CPU is now the binding constraint. FINDINGS 28. Three things that were believed and are not true: - The dual-display-path plan of FINDINGS 24.5/25.6 is incoherent. The compose path needs a RAM copy of the previous reconstruction; the direct path's selling point is that it keeps none. Mixing them shows stale pixels on 70 of 120 frames, worst frame 18.8% of the screen. Every coherent repair is dearer than not mixing, and 24.5's two figures were both copies with no decode in either, so there was never a crossover to find. One path ships, and the 96KB reference frame is gone. tools/analysis/10_pathmix_drift.py keeps the counterexample runnable; check.sh asserts it still reproduces. - The four block modes do not cost the same. V1 300, V4 448, RAW 400 cycles against the old model's flat 207.8. V4 is 25% of blocks and 50% of the cycles, and the mode decision charges it bytes it does not charge cycles for. tools/analysis/11_cpu_budget.py reproduces all four frames timed on the 68000 to within 1 point. Hand-derived timings agree to 0.5% on V1. - The container is big-endian but not aligned. Variable-length records laid end to end put frame 1's length field at an odd address, and move.l (a0)+ there is an address error: frame 0 decoded perfectly and then vectored into the IPL for 59 emulated seconds looking like a hang. Found by dumping PC, not by reading the source. Also: an all-V1 frame, the cheapest possible full redraw, is 110.5% of budget. No mode assignment fits a scene cut at 12fps. That one needs a decision, not a measurement. Next: charge cycles in the mode decision and bisect against 833,333 per frame, the way session 6 bisects lam against bytes -- but with no bucket, because a late frame cannot be banked. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -0,0 +1,136 @@
|
||||
#!/usr/bin/env python3
|
||||
"""REGRESSION TEST for the dual-display-path coherency defect (FINDINGS 28).
|
||||
|
||||
FAILS ON PURPOSE for the player design FINDINGS 24.5/25.6 specified, which is
|
||||
why it exists: it is the counterexample, kept runnable. Pass `--fix <strategy>`
|
||||
to check that a proposed player is coherent instead.
|
||||
|
||||
The defect. Two display paths were specified and priced, and the plan was to
|
||||
pick between them per frame on the non-SKIP block count:
|
||||
|
||||
compose-in-RAM-then-blit flat 53.6% of a 12fps frame
|
||||
decode-direct-to-GVRAM 76.6% x (non-SKIP fraction), and -- quoting
|
||||
FINDINGS 24.5 -- "no RAM reference frame is needed"
|
||||
|
||||
Both statements are true in isolation and incompatible together. The compose
|
||||
path's whole job is to assemble a FULL frame in RAM so the blit can be
|
||||
row-linear, and the pixels it does not decode this frame (the SKIP blocks) can
|
||||
only come from a RAM copy of the previous reconstruction. The direct path
|
||||
deliberately never writes that copy. So every direct frame silently invalidates
|
||||
the reference the next compose frame reads, and the stale pixels go to screen.
|
||||
|
||||
This is FINDINGS 26 again in different clothing: two code paths that disagree
|
||||
about what "the previous frame" means. 26 was caught between encoder and
|
||||
decoder; this one is between a decoder and itself.
|
||||
|
||||
Strategies (--fix):
|
||||
none the specified player: direct writes GVRAM only. UNSOUND
|
||||
dual direct also writes the RAM reference (costs ~1.52x). sound
|
||||
resync on direct->compose, re-read GVRAM into RAM first. sound
|
||||
compose never use the direct path. sound
|
||||
direct never use the compose path. sound
|
||||
|
||||
Needs a container: defaults to the worst sustained window on the disc at the
|
||||
shipping profile (docs/STATUS.md, reproducing the rate-control result). Costs
|
||||
no encode -- it reads the emitted bitstream, ~3 s.
|
||||
"""
|
||||
import sys, os, argparse
|
||||
sys.path.insert(0, "tools/encoder")
|
||||
import numpy as np
|
||||
from dlx import DLX
|
||||
|
||||
# FINDINGS 24: measured on the emulated 68000, instruction cycles only.
|
||||
BLIT_PCT, DIRECT_PCT = 53.6, 76.6
|
||||
CROSSOVER = 100 * BLIT_PCT / DIRECT_PCT
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("container", nargs="?",
|
||||
default="tmp/rc_fr_singe_sasi_rcprofile.dlx")
|
||||
ap.add_argument("--fix", default="none",
|
||||
choices=("none", "dual", "resync", "compose", "direct"))
|
||||
a = ap.parse_args()
|
||||
|
||||
if not os.path.exists(a.container):
|
||||
sys.exit(f"missing {a.container} -- see docs/STATUS.md, "
|
||||
f"'Reproducing the rate-control result'")
|
||||
|
||||
d = DLX(a.container)
|
||||
print(f"{a.container}: {d.W}x{d.H} {d.nframes} frames, {d.nb} blocks/frame, "
|
||||
f"k1={d.k1} k4={d.k4}")
|
||||
print(f"path choice: compose if non-SKIP > {CROSSOVER:.1f}% of blocks "
|
||||
f"(53.6% flat vs 76.6% x fraction), strategy={a.fix}\n")
|
||||
|
||||
# gv = what is on screen. ram = the player's RAM reference frame.
|
||||
# truth = what a correct player displays. All are palette-index canvases.
|
||||
gv = np.zeros((d.H, d.W), np.uint8)
|
||||
ram = np.zeros((d.H, d.W), np.uint8)
|
||||
truth = np.zeros((d.H, d.W), np.uint8)
|
||||
|
||||
def put(canvas, blks):
|
||||
for i, blk in blks.items():
|
||||
by, bx = divmod(i, d.nbx)
|
||||
canvas[by*4:by*4+4, bx*4:bx*4+4] = blk
|
||||
|
||||
drift_px, used, switches, resyncs = [], [], 0, 0
|
||||
prev_path = None
|
||||
for f in range(d.nframes):
|
||||
mode, blks = d.blocks(f)
|
||||
put(truth, blks)
|
||||
|
||||
frac = 100 * (mode != 0).mean()
|
||||
if a.fix == "compose": path = "compose"
|
||||
elif a.fix == "direct": path = "direct"
|
||||
else: path = "compose" if frac > CROSSOVER else "direct"
|
||||
|
||||
if path == "compose":
|
||||
if a.fix == "resync" and prev_path == "direct":
|
||||
ram = gv.copy() # re-read GVRAM into the RAM reference
|
||||
resyncs += 1
|
||||
put(ram, blks)
|
||||
gv = ram.copy() # full row-linear blit
|
||||
else:
|
||||
put(gv, blks)
|
||||
if a.fix == "dual":
|
||||
put(ram, blks) # keep the reference coherent as we go
|
||||
|
||||
if prev_path is not None and path != prev_path:
|
||||
switches += 1
|
||||
prev_path = path
|
||||
used.append(path)
|
||||
drift_px.append(int((gv != truth).sum()))
|
||||
|
||||
drift_px = np.array(drift_px)
|
||||
npx = d.H * d.W
|
||||
nc = used.count("compose")
|
||||
print(f"path used: compose {nc}/{d.nframes} ({100*nc/d.nframes:.0f}%), "
|
||||
f"direct {d.nframes-nc} -- {switches} switches between them"
|
||||
+ (f", {resyncs} resyncs" if resyncs else ""))
|
||||
bad = int((drift_px > 0).sum())
|
||||
print(f"\nframes displaying pixels no correct player would display: "
|
||||
f"{bad}/{d.nframes}")
|
||||
if bad:
|
||||
print(f" worst frame {drift_px.max()} px "
|
||||
f"({100*drift_px.max()/npx:.1f}% of the screen), "
|
||||
f"mean {drift_px.mean():.0f} px ({100*drift_px.mean()/npx:.1f}%)")
|
||||
first = int(np.argmax(drift_px > 0))
|
||||
print(f" first corrupt frame: {first} (path={used[first]}, "
|
||||
f"previous={used[first-1] if first else '-'})")
|
||||
|
||||
# Cost of the strategy, in % of a 12fps frame budget. 'dual' pays 1.52x on the
|
||||
# direct path: the same block written twice, +108 cycles on 208 (FINDINGS 28.2).
|
||||
mult = 1.52 if a.fix == "dual" else 1.0
|
||||
cost = np.array([BLIT_PCT if p == "compose" else DIRECT_PCT * mult *
|
||||
(d.modes(f) != 0).mean()
|
||||
for f, p in enumerate(used)])
|
||||
if a.fix == "resync":
|
||||
for f in range(1, d.nframes):
|
||||
if used[f] == "compose" and used[f-1] == "direct":
|
||||
cost[f] += BLIT_PCT # the GVRAM->RAM re-read is a full frame
|
||||
print(f"\ndisplay cost: median {np.median(cost):.1f}% "
|
||||
f"p90 {np.percentile(cost,90):.1f}% max {cost.max():.1f}% "
|
||||
f"of a 12fps frame budget")
|
||||
over = int((cost > 100).sum())
|
||||
if over:
|
||||
print(f" frames that do NOT fit in the budget at all: {over}/{d.nframes}")
|
||||
|
||||
sys.exit(1 if bad else 0)
|
||||
Reference in New Issue
Block a user