The bus is 4x idle while the CPU is pinned: price the trade
The codec was designed when bytes were scarce, so every decision in it trades
cycles to save bytes. That is now backwards: sasi spends 110 KB/s of a 488 KB/s
pipe while missing 31% of frames on CPU.
The cheapest thing a 68000 can be handed is the most expensive thing to store.
Measured, per pixel: row-linear copy from word-expanded memory 9.08 cycles,
block-order 12.98, V1 codebook 18.74, RAW byte literals 25.03. So the 1024-byte
stride costs 43% and unpacking bytes to words costs more than the write itself.
Pricing one new mode -- a per-row span of word-expanded literals movem.l'd
straight from the stream buffer -- against the UNCHANGED mode maps:
sasi median 74.4% -> 43.0%, worst 136.2% -> 106.2%, misses 37 -> 8/120,
101.7 -> 453.2 KB/s
scsi median 94.9% -> 69.4%, misses 51 -> 18/120, 272 -> 479.7 KB/s
scsi gains less precisely because it has less idle bandwidth left to trade.
Two consequences worth flagging. A word-expanded literal block derives to ~240
cycles, cheaper than V1's measured 299.9 and pixel-exact -- so every codebook
mode is CPU-dominated by a literal, and the codebook is a byte optimisation
that now costs cycles. And 28.5's "a scene cut cannot fit at 12fps" reopens:
CPU needs >=19% of the frame as spans, the bus allows <=39%, and that interval
is not empty.
DERIVED, NOT MEASURED, and labelled as such everywhere. The 9.08 cycles/pixel
is real but was measured at full row width with 12-register bursts, so short
spans are flattered. Measuring one span on the 68000 is now step 0 of the next
session, ahead of the cost-aware mode decision, because it changes the mode set
that decision optimises over.
FINDINGS 29. tools/analysis/12_span_tradeoff.py.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -0,0 +1,111 @@
|
||||
#!/usr/bin/env python3
|
||||
"""What does spending the idle bus bandwidth buy back in CPU cycles?
|
||||
|
||||
python3 tools/analysis/12_span_tradeoff.py [container.dlx] [--bus 488]
|
||||
|
||||
FINDINGS 28 leaves the decoder CPU-bound at 110 KB/s on a 488 KB/s pipe. Every
|
||||
codec decision was made when bytes were scarce, so each one trades cycles to
|
||||
save them -- and the cheapest thing a 68000 can be handed is the most expensive
|
||||
thing to store: word-expanded pixels in row-linear runs.
|
||||
|
||||
This prices ONE new mode against the real mode maps: a per-row SPAN of
|
||||
word-expanded literals, `movem.l`-ed straight from the stream buffer into GVRAM.
|
||||
A run of L horizontally adjacent dirty blocks becomes 4 spans of 4L pixels.
|
||||
|
||||
DERIVED, NOT MEASURED (FINDINGS 29). The 9.08 cycles/pixel is measured
|
||||
(FINDINGS 24 V1) but at full row width with 12-register bursts; SPAN_OVERHEAD is
|
||||
hand-derived. Short spans are therefore flattered. Measure before believing --
|
||||
FINDINGS 29.5 item 1.
|
||||
|
||||
The mode maps are NOT re-optimised: this only re-codes regions the encoder
|
||||
already chose to redraw, so it is a lower bound on what a cost-aware encoder
|
||||
would find.
|
||||
"""
|
||||
import sys, os, argparse
|
||||
sys.path.insert(0, "tools/encoder")
|
||||
import numpy as np
|
||||
from dlx import DLX
|
||||
|
||||
FRAME_CYC = 833333.0 # 12fps at 10 MHz
|
||||
AUDIO_KBPS = 7.8
|
||||
|
||||
CYC_PX_ROWLIN = 446286 / 49152. # 9.08, FINDINGS 24 V1 (measured)
|
||||
C_V1, C_V4, C_RAW = 299.9, 448.2, 400.4 # FINDINGS 28.2 (measured)
|
||||
C_SKIP_CLUSTERED, C_SKIP_MIXED = 13.25, 45.0
|
||||
SPAN_OVERHEAD = 50.0 # per span, DERIVED
|
||||
SPAN_BYTES_PX = 2 # word-expanded: 1 pixel = 1 word
|
||||
SPAN_HDR = 3 # x, count, and a byte of slack
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("container", nargs="?",
|
||||
default="tmp/rc_fr_singe_sasi_rcprofile.dlx")
|
||||
ap.add_argument("--bus", type=float, default=488.0,
|
||||
help="sustained KB/s the pipe delivers (FINDINGS 21)")
|
||||
ap.add_argument("--fps", type=float, default=12.0)
|
||||
a = ap.parse_args()
|
||||
if not os.path.exists(a.container):
|
||||
sys.exit(f"missing {a.container}")
|
||||
|
||||
BYTE_BUD = (a.bus - AUDIO_KBPS) * 1024 / a.fps
|
||||
d = DLX(a.container)
|
||||
BLK_C = {1: C_V1, 2: C_V4, 3: C_RAW}
|
||||
BLK_B = {1: 1, 2: 4, 3: 16}
|
||||
|
||||
rows = []
|
||||
for f in range(d.nframes):
|
||||
mode = d.modes(f)
|
||||
g = mode.reshape(-1, 4)
|
||||
allskip = (g == 0).all(1)
|
||||
base = allskip.sum() * 4 * C_SKIP_CLUSTERED
|
||||
mm = g[~allskip]
|
||||
base += (mm == 0).sum() * C_SKIP_MIXED
|
||||
for k, c in BLK_C.items():
|
||||
base += (mm == k).sum() * c
|
||||
base_b = d.mode_bytes + sum(BLK_B.get(int(x), 0) for x in mode)
|
||||
|
||||
m = mode.reshape(d.nby, d.nbx)
|
||||
cand = []
|
||||
for by in range(d.nby):
|
||||
dirty = m[by] != 0
|
||||
i = 0
|
||||
while i < d.nbx:
|
||||
if not dirty[i]:
|
||||
i += 1
|
||||
continue
|
||||
j = i
|
||||
while j < d.nbx and dirty[j]:
|
||||
j += 1
|
||||
L = j - i
|
||||
cur_c = sum(BLK_C[int(b)] for b in m[by][i:j])
|
||||
cur_b = sum(BLK_B[int(b)] for b in m[by][i:j])
|
||||
span_c = 4 * (SPAN_OVERHEAD + 4 * L * CYC_PX_ROWLIN)
|
||||
span_b = 4 * (SPAN_HDR + 4 * L * SPAN_BYTES_PX)
|
||||
if span_c < cur_c:
|
||||
cand.append((cur_c - span_c, span_b - cur_b, L))
|
||||
i = j
|
||||
|
||||
cand.sort(key=lambda s: -(s[0] / max(s[1], 1))) # best cycles per byte
|
||||
cyc, byt, taken = base, base_b, 0
|
||||
for dc, db, L in cand:
|
||||
if byt + db <= BYTE_BUD:
|
||||
cyc -= dc; byt += db; taken += 1
|
||||
rows.append((base, cyc, base_b, byt, len(cand), taken))
|
||||
|
||||
base, new, bb, nb, ncand, ntaken = map(np.array, list(zip(*rows)))
|
||||
pc = lambda v: 100 * v / FRAME_CYC
|
||||
|
||||
print(f"{a.container}: {d.nframes} frames")
|
||||
print(f"bus {a.bus:.0f} KB/s - {AUDIO_KBPS} audio -> {BYTE_BUD:,.0f} B/frame "
|
||||
f"at {a.fps:g}fps\n")
|
||||
print(f"{'':<26}{'today':>12}{'+ literal spans':>18}")
|
||||
for label, fn in (("median frame", np.median),
|
||||
("p90 frame", lambda v: np.percentile(v, 90)),
|
||||
("worst frame", np.max)):
|
||||
print(f" {label:<24}{pc(fn(base)):>11.1f}%{pc(fn(new)):>17.1f}%")
|
||||
print(f" {'frames missing budget':<24}{int((base>FRAME_CYC).sum()):>8}/{d.nframes}"
|
||||
f"{int((new>FRAME_CYC).sum()):>14}/{d.nframes}")
|
||||
print(f" {'bitrate':<24}{bb.mean()*a.fps/1024:>10.1f} KB/s"
|
||||
f"{nb.mean()*a.fps/1024:>13.1f} KB/s")
|
||||
print(f"\nspans taken: {ntaken.sum()} of {ncand.sum()} candidate runs "
|
||||
f"({100*ntaken.sum()/max(ncand.sum(),1):.0f}%) -- the rest priced out by the bus")
|
||||
print("\nDERIVED, NOT MEASURED: see FINDINGS 29.5 before acting on this.")
|
||||
Reference in New Issue
Block a user