Files
Dragon-s-Lair-X68k/tools/encoder/profile_gen.py
T
prosolis 48e912de8b Size against 4 Mbps: peaks break the scsi profile; DMA steal is not free
User clarified the bandwidth figure is 4 Mbps (488 KB/s), not 4 MB/s -- ~8x
tighter than the previous commit reasoned against. Two consequences, plus a
correction to session 1.

1. The scsi profile committed in f0f2f80 DOES NOT FIT. Its mean is a
   comfortable 52% of the pipe but it PEAKS at 96.4% (470.8 KB/s on scene
   00020), and a frame that arrives late is a dropped frame, not a slow one.
   Peak/mean is 1.4-1.9x even on 1.2-1.7s clips. Sizing a real-time stream on
   the mean was the error. Flagged in STATUS rather than silently retuned,
   because the fix is rate control, not a lower lam.

   This promotes ratectl.py -- written in session 2, never wired into
   encode.py -- from a loose end to the highest-value work in the repo. It is
   worth a full step on the quality ladder (lam=25 -> lam=10, +0.7/+1.2 dB)
   because it allows sizing for the mean instead of the peak.

2. Pixel-exact is off the table at this bandwidth: lam=0 needs 92-97% of the
   pipe. The previous commit's "if SCSI sustains >=800 KB/s, ship transparent"
   conclusion only applies at roughly double the user's figure.

3. FINDINGS 5 said that because transfers are DMA, streaming "costs essentially
   no CPU" and the 68000 is "nearly idle". That is wrong. The HD63450 steals
   ~8 clocks per 16-bit word: 10-20% of the machine at the rates the profiles
   now use, on top of a 38% full-frame blit. Bandwidth and CPU are one budget.

Adds tools/encoder/profile_gen.py, which derives lam FROM a bandwidth figure
(accounting for audio, peak/mean and DMA steal) instead of reading it off the
knee of the RD curve, and docs/BENCHMARK.md covering how to actually measure
the subsystem -- including why MAME cannot answer the bandwidth question and
would be the same class of error as the FINDINGS 4 traps.

The 4 Mbps figure is user-supplied and its provenance is not recorded; every
profile now hangs off it.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 12:12:34 -07:00

118 lines
5.6 KiB
Python

#!/usr/bin/env python3
"""Derive quality profiles FROM a measured bandwidth, instead of guessing lam.
python3 tools/encoder/profile_gen.py --bw-kbps 488 --name scsi
python3 tools/encoder/profile_gen.py --bw-mbps 4 # same thing
Session 2 set the profile bitrates by eye off the rate-distortion knee, which
was wrong twice over (FINDINGS 17.1). This inverts the dependency: give it a
bandwidth and it returns the lam that fits, with the headroom accounted for.
Three things eat the pipe before video gets any:
1. AUDIO -- 7.8 KB/s of MSM6258 ADPCM, constant.
2. PEAK/MEAN -- measured 1.4-1.9x (FINDINGS 18). The disk delivers a
SUSTAINED rate; a frame that overruns is a DROPPED frame.
Either size for the peak, or rate-control to the mean and
carry a bucket. We do the latter, so we need bucket depth
rather than peak headroom -- but until rate control is
actually wired in (it is not), size for the peak.
3. DMA CYCLE-STEAL -- the HD63450 steals ~8 clocks per 16-bit word from the
68000. At 488 KB/s that is 20% of the CPU, on top of the
blit. Bandwidth and CPU are NOT independent budgets.
FINDINGS 5 said streaming "costs essentially no CPU";
that is wrong -- cycle-stealing DMA is not free DMA.
The rate-distortion points are MEASURED (FINDINGS 17.4), not modelled, so this
interpolates real data rather than fitting a curve to a guess.
"""
import argparse
AUDIO_KBPS = 7.8
CLK = 10_000_000
FPS = 12
BLIT_FULL_FRAME_PCT = 38.3 # FINDINGS 17.2
DMA_CLOCKS_PER_WORD = 8 # FINDINGS 5 (ESTIMATE, from HD63450 timing)
# (lam, KB/s, PSNR) measured on the two probe scenes -- FINDINGS 17.4.
# 00146 is the harder scene; we size against it so profiles are not tuned to
# the easy case. Rates are RAW payload: entropy coding is ruled out (17.2).
CURVE = [
# lam 00020 KB/s 00020 dB 00146 KB/s 00146 dB
( 0, 442.1, 39.90, 467.6, 35.25),
( 10, 248.1, 39.38, 305.2, 32.27),
( 25, 182.2, 38.68, 193.5, 31.04),
( 60, 108.0, 36.94, 103.1, 29.61),
( 150, 55.6, 35.31, 56.1, 28.63),
( 300, 44.1, 34.80, 44.4, 28.28),
( 800, 32.5, 33.87, 36.1, 27.77),
]
CEILING = {"00020": 39.90, "00146": 35.25}
PEAK_OVER_MEAN = 1.9 # measured worst case, FINDINGS 18
def dma_steal_pct(kbps):
return (kbps * 1024 / 2) * DMA_CLOCKS_PER_WORD / CLK * 100
def pick(bw_kbps, peak_factor=PEAK_OVER_MEAN, margin=0.85, rate_controlled=False):
"""Largest-quality lam whose worst-case demand fits inside bw_kbps."""
usable = bw_kbps * margin - AUDIO_KBPS
factor = 1.0 if rate_controlled else peak_factor
allow_mean = usable / factor
for lam, k20, d20, k146, d146 in CURVE:
worst = max(k20, k146)
if worst <= allow_mean:
return dict(lam=lam, mean_kbps=worst, peak_kbps=worst * factor,
psnr20=d20, psnr146=d146,
loss20=CEILING["00020"] - d20,
loss146=CEILING["00146"] - d146,
allow_mean=allow_mean, usable=usable)
return None
def main():
ap = argparse.ArgumentParser()
g = ap.add_mutually_exclusive_group(required=True)
g.add_argument("--bw-kbps", type=float)
g.add_argument("--bw-mbps", type=float, help="megaBITS/sec")
ap.add_argument("--name", default="profile")
ap.add_argument("--margin", type=float, default=0.85,
help="fraction of the pipe we allow ourselves (seeks, "
"container overhead, and the fact that the bandwidth "
"figure itself is folklore)")
ap.add_argument("--rate-controlled", action="store_true",
help="assume the leaky bucket absorbs peaks (NOT YET TRUE "
"-- ratectl.py is written but not wired into encode.py)")
a = ap.parse_args()
bw = a.bw_kbps if a.bw_kbps else a.bw_mbps * 1_000_000 / 8 / 1024
src = f"{a.bw_mbps} Mbps" if a.bw_mbps else f"{a.bw_kbps} KB/s"
print(f"bandwidth {src} = {bw:.0f} KB/s sustained")
print(f" usable at {a.margin:.0%} margin : {bw*a.margin:.0f} KB/s")
print(f" less audio ({AUDIO_KBPS}) : {bw*a.margin-AUDIO_KBPS:.0f} KB/s for video")
if not a.rate_controlled:
print(f" less peak/mean {PEAK_OVER_MEAN}x : "
f"{(bw*a.margin-AUDIO_KBPS)/PEAK_OVER_MEAN:.0f} KB/s mean allowance")
else:
print(" peaks absorbed by rate control (bucket depth must be validated)")
r = pick(bw, margin=a.margin, rate_controlled=a.rate_controlled)
if r is None:
print("\n NO PROFILE FITS -- even lam=800 overruns. Lower the framerate,")
print(" the resolution, or get more bandwidth.")
return
steal = dma_steal_pct(r["peak_kbps"])
print(f"\n -> {a.name}: lam={r['lam']}, {r['mean_kbps']:.0f} KB/s mean, "
f"{r['peak_kbps']:.0f} KB/s peak")
print(f" quality 00020 {r['psnr20']:.2f} dB (-{r['loss20']:.2f} from ceiling)")
print(f" 00146 {r['psnr146']:.2f} dB (-{r['loss146']:.2f} from ceiling)")
print(f" CPU blit {BLIT_FULL_FRAME_PCT:.0f}% + DMA steal {steal:.1f}% "
f"= {BLIT_FULL_FRAME_PCT+steal:.0f}% of the frame budget")
if BLIT_FULL_FRAME_PCT + steal > 85:
print(" WARNING: CPU is now the binding constraint, not the bus.")
if __name__ == "__main__":
main()