USER DECISION: drop the `sasi` profile. Not on bandwidth -- on capacity. A SASI volume is 40 MB, and the 22.8 min of unique scene footage on the source Blu-ray (streams 00000-00201, measured, not recalled) is 146 MiB at the LOWEST rate this codec makes -- more than the machine's whole 4-unit SASI space. `scsi` is the only profile now. FINDINGS 32. Then the user asked whether we were drawing the wrong conclusions about PIO vs DMA, and we were, more broadly than the question implied. Every CPU figure in FINDINGS 24-34 is scored against the full 833,333 cycles/frame with nothing subtracted for moving the bitstream off disk. Debiting the HD63450 cycle-steal at the long-standing 8 clk/word ESTIMATE, "1 frame of 120 misses" becomes 84 of 120, median 112.4%. PIO at the span rate is 99.8% of the machine. Spans buy cycles by spending bandwidth and the bandwidth returns as steal, so 31.6's "fits completely" becomes a worst frame of 114.3%. 10 fps absorbs it: median 93.7%, 1/120. FINDINGS 35. `11_cpu_budget.py` takes --io dma|pio|none, defaults to dma, and warns if asked for none. Also landed: - item 1 done: the cost model checked against the 68000 on a cost-aware container, -3.07% to +0.01%, whole-window mean -1.22%. FINDINGS 34. - item 4 done: the container carries its own 4-byte record alignment (DLX2). 94/120 record starts were on odd addresses -- an address error, not a slow read -- now 0/120 for 16 B/s. Re-encoding reproduces 31.1 exactly. FINDINGS 33. - a `scsi` window does not fit the 2 MB machine the rig emulates (2.84 MB of stream past a 0x200000 ceiling). The gate now verifies 80 of 120 frames and SAYS so, and fails loudly when the pass does not complete, instead of reporting a phantom 49,005-pixel diff. FINDINGS 36. Three near-misses this session had one shape: an unobservable run nearly produced a false finding. stdbuf -oL on any MAME job that prints progress -- a file is block-buffered too, and a run that is merely finishing looks exactly like one that is wedged. check.sh ALL GREEN. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
160 lines
8.0 KiB
Python
160 lines
8.0 KiB
Python
#!/usr/bin/env python3
|
|
"""Per-frame CPU cost of the real decoder, from MEASURED per-mode block costs.
|
|
|
|
python3 tools/analysis/11_cpu_budget.py [container.dlx]
|
|
|
|
FINDINGS 24.5 priced the display path as "76.6% of a 12fps frame x the non-SKIP
|
|
block fraction", i.e. every non-SKIP block costs the same. It does not: the four
|
|
block modes were measured separately on the 68000 (synthetic single-mode frames,
|
|
tools/bench/prep_dlx.py) and V4 costs 1.5x V1. Since V4 is roughly half of all
|
|
non-SKIP blocks on hard content, the old model runs ~1.8x optimistic exactly
|
|
where it matters.
|
|
|
|
This applies the measured costs to a real container's mode histograms and
|
|
reports what fraction of frames actually fit 833,333 cycles.
|
|
|
|
Costs are MEASURED (tools/bench/decode.lua), cross-checked against hand-derived
|
|
MC68000 timings in FINDINGS 28.4. They are instruction cycles against
|
|
zero-wait-state memory, so like every figure in this project since FINDINGS 24
|
|
they are a LOWER BOUND -- real GVRAM stalls the CPU.
|
|
"""
|
|
import sys, os, argparse
|
|
sys.path.insert(0, "tools/encoder")
|
|
import numpy as np
|
|
from dlx import DLX
|
|
import vq_hybrid as H
|
|
import ratectl as RC
|
|
RC_AUDIO_BPS = RC.AUDIO_KBPS * 1024
|
|
|
|
# Machine clocks, confirmed from MAME 0.277 src/mame/sharp/x68k.cpp:1133/1194/
|
|
# 1200 -- not recalled. x68000 and x68ksupr are BOTH 40_MHz_XTAL/4 = 10 MHz;
|
|
# only the XVI is faster, at 33.33_MHz_XTAL/2. So "has SCSI" and "has a faster
|
|
# CPU" are different sets of machines: the Super has SCSI at 10 MHz.
|
|
CLOCKS = {"stock": 10.0, "super": 10.0, "xvi": 33.33 / 2, "x68030": 25.0}
|
|
FPS = 12
|
|
|
|
# Cycles per block, measured on the emulated 68000 (synthetic single-mode
|
|
# frames). Defined in tools/encoder/vq_hybrid.py, which is where the mode
|
|
# decision needs them too -- one copy, not two, so a re-measurement cannot
|
|
# leave the encoder and the scorer disagreeing.
|
|
C_V1, C_V4, C_RAW = H.C_V1, H.C_V4, H.C_RAW
|
|
C_SKIP_FAST, C_SKIP_MIXED = H.C_SKIP_CLUSTERED, H.C_SKIP_MIXED
|
|
cycles = H.cycles
|
|
|
|
ap = argparse.ArgumentParser()
|
|
ap.add_argument("container", nargs="?",
|
|
default="tmp/rc_fr_singe_sasi_rcprofile.dlx")
|
|
ap.add_argument("--machine", default="stock", choices=list(CLOCKS),
|
|
help="which X68000's clock to budget against (default stock)")
|
|
ap.add_argument("--fps", type=float, default=FPS)
|
|
# FINDINGS 35: the frame budget has never had the disk in it. The bitstream has
|
|
# to be moved off SCSI into the ring buffer, and on this machine that costs CPU
|
|
# whether it is DMA (the HD63450 cycle-steals) or PIO (the 68000 moves every
|
|
# byte). Default ON, because scoring a decoder against a budget that assumes the
|
|
# data arrives for free is exactly the mistake 35 was raised to stop.
|
|
ap.add_argument("--io", default="dma", choices=["dma", "pio", "none"],
|
|
help="how the bitstream reaches RAM (default dma)")
|
|
ap.add_argument("--dma-clocks-per-word", type=float, default=8.0,
|
|
help="HD63450 cycle-steal. ESTIMATE from FINDINGS 5, NEVER "
|
|
"MEASURED, and the most load-bearing unmeasured number "
|
|
"in the project (FINDINGS 35.3)")
|
|
ap.add_argument("--pio-clocks-per-byte", type=float, default=12.0,
|
|
help="hand-derived floor for a 68000 register-to-RAM copy")
|
|
a = ap.parse_args()
|
|
CPUHZ = CLOCKS[a.machine] * 1e6
|
|
FPS = a.fps
|
|
FRAME = CPUHZ / FPS
|
|
if not os.path.exists(a.container):
|
|
sys.exit(f"missing {a.container}")
|
|
|
|
d = DLX(a.container)
|
|
|
|
# --- what the transfer costs, from the container's own byte rate
|
|
vid_bps = sum(n + 4 for (_, n) in d.frames) / d.nframes * d.fps
|
|
io_bps = vid_bps + RC_AUDIO_BPS
|
|
if a.io == "dma":
|
|
io_cycles_per_s = (io_bps / 2) * a.dma_clocks_per_word
|
|
elif a.io == "pio":
|
|
io_cycles_per_s = io_bps * a.pio_clocks_per_byte
|
|
else:
|
|
io_cycles_per_s = 0.0
|
|
io_pct = 100 * io_cycles_per_s / CPUHZ
|
|
FRAME_NET = FRAME * (1 - io_pct / 100)
|
|
|
|
modes = [d.modes(f) for f in range(d.nframes)]
|
|
cyc = np.array([cycles(m) for m in modes])
|
|
pct = 100 * cyc / FRAME_NET
|
|
ns = np.array([100 * (m != 0).mean() for m in modes])
|
|
|
|
print(f"{a.container}: {d.nframes} frames, {d.nb} blocks/frame")
|
|
print(f"budget: {a.machine} @ {CLOCKS[a.machine]:.2f} MHz, {FPS:g} fps "
|
|
f"-> {FRAME:,.0f} cycles/frame")
|
|
print(f" I/O ({a.io}): {io_bps/1024:.1f} KB/s costs {io_pct:.1f}% of the CPU "
|
|
f"-> {FRAME_NET:,.0f} cycles/frame left for decoding")
|
|
if a.io == "dma":
|
|
print(f" {a.dma_clocks_per_word:g} clocks/word is an ESTIMATE (FINDINGS 5), "
|
|
f"never measured -- see FINDINGS 35.3")
|
|
elif a.io == "none":
|
|
print(" WARNING: --io none scores the decoder as if the disk were free. "
|
|
"That is the\n premise FINDINGS 35 overturned; every 'N frames miss' "
|
|
"figure before session 9\n was computed this way.")
|
|
if a.machine != "stock":
|
|
print(" (derived: scaled by clock from cycles measured on the 10 MHz core.\n"
|
|
" MAME 0.277 marks x68ksupr/x68kxvi/x68030 MACHINE_NOT_WORKING, so\n"
|
|
" this is not measured on those machines and ignores any difference\n"
|
|
" in memory timing.)")
|
|
print(f"measured block costs: SKIP {C_SKIP_FAST*4:.0f}/4 (clustered) "
|
|
f"{C_SKIP_MIXED:.0f} (mixed) V1 {C_V1:.0f} V4 {C_V4:.0f} RAW {C_RAW:.0f} cycles\n")
|
|
|
|
# --- validation against the four real frames timed on the 68000. These
|
|
# timings belong to ONE container; quoting them against any other would be
|
|
# comparing a model of this stream to a measurement of a different one.
|
|
TIMED = "tmp/rc_fr_singe_sasi_rcprofile.dlx"
|
|
TIMED_FRAMES = (("min non-SKIP", 15.4, 31.5), ("median", 48.1, 73.8),
|
|
("p90", 82.5, 116.4), ("max non-SKIP", 100.0, 135.8))
|
|
if (os.path.abspath(a.container) == os.path.abspath(TIMED)
|
|
and a.machine == "stock" and a.fps == 12):
|
|
print("model vs the frames actually timed on the 68000:")
|
|
for label, frac, meas in TIMED_FRAMES:
|
|
i = int(np.argmin(abs(ns - frac)))
|
|
print(f" {label:<14} non-SKIP {ns[i]:5.1f}% model {pct[i]:6.1f}% "
|
|
f"measured {meas:5.1f}% error {pct[i]-meas:+.1f} pt")
|
|
else:
|
|
print(f"(no 68000 timings for this container/machine -- the model was "
|
|
f"validated to\n within 1 pt on {TIMED} at stock/12fps;\n"
|
|
f" run tools/bench/decode.lua to time another container)")
|
|
|
|
print(f"\nper-frame cost, % of a {FPS:g}fps frame budget:")
|
|
print(f" measured-cost model: median {np.median(pct):5.1f} "
|
|
f"p90 {np.percentile(pct,90):5.1f} max {pct.max():5.1f}")
|
|
if a.machine == "stock" and a.fps == 12:
|
|
# 24.5's 76.6% is a 10 MHz / 12 fps figure; quoting it at another clock or
|
|
# framerate would be comparing against a model that was never stated there.
|
|
old = 76.6 * ns / 100
|
|
print(f" FINDINGS 24.5 model: median {np.median(old):5.1f} "
|
|
f"p90 {np.percentile(old,90):5.1f} max {old.max():5.1f} "
|
|
f"(optimistic by {np.median(pct)/np.median(old):.2f}x at the median)")
|
|
|
|
miss = pct > 100
|
|
print(f"\nframes that do NOT fit {FRAME_NET:,.0f} cycles: {miss.sum()}/{d.nframes} "
|
|
f"({100*miss.mean():.0f}%)")
|
|
print(f" sustainable framerate if EVERY frame must fit: "
|
|
f"{CPUHZ*(1-io_pct/100)/cyc.max():.1f} fps; at the mean frame "
|
|
f"{CPUHZ*(1-io_pct/100)/cyc.mean():.1f} fps")
|
|
if miss.any():
|
|
print(f" worst {pct.max():.1f}% -- {(pct.max()-100)/100*1000/FPS:.0f} ms late "
|
|
f"on an {1000/FPS:.0f} ms frame")
|
|
print(f" the budget is first missed at {ns[miss].min():.1f}% non-SKIP blocks")
|
|
|
|
# Where do the cycles go? This is what a cost-aware mode decision would act on.
|
|
tot = np.array([[(m == k).sum() for k in range(4)] for m in modes]).sum(0)
|
|
spend = tot * np.array([C_SKIP_MIXED, C_V1, C_V4, C_RAW])
|
|
print(f"\nwhere the cycles go, over the whole window:")
|
|
for k, n in enumerate(("SKIP", "V1", "V4", "RAW")):
|
|
print(f" {n:<5} {100*tot[k]/tot.sum():5.1f}% of blocks "
|
|
f"{100*spend[k]/spend.sum():5.1f}% of the cycles")
|
|
print(f"\nV4 is {C_V4/C_V1:.2f}x a V1 block for {4}x the payload bytes. Since "
|
|
f"session 8 the mode\ndecision charges it BOTH (decide(ctx, lam, mu), "
|
|
f"FINDINGS 31), which is why V4 is now\nthe rarest non-SKIP mode here -- "
|
|
f"a byte-rich profile buys its way out to RAW instead.")
|