Files
Dragon-s-Lair-X68k/docs/BENCHMARK.md
T
prosolis 48e912de8b Size against 4 Mbps: peaks break the scsi profile; DMA steal is not free
User clarified the bandwidth figure is 4 Mbps (488 KB/s), not 4 MB/s -- ~8x
tighter than the previous commit reasoned against. Two consequences, plus a
correction to session 1.

1. The scsi profile committed in f0f2f80 DOES NOT FIT. Its mean is a
   comfortable 52% of the pipe but it PEAKS at 96.4% (470.8 KB/s on scene
   00020), and a frame that arrives late is a dropped frame, not a slow one.
   Peak/mean is 1.4-1.9x even on 1.2-1.7s clips. Sizing a real-time stream on
   the mean was the error. Flagged in STATUS rather than silently retuned,
   because the fix is rate control, not a lower lam.

   This promotes ratectl.py -- written in session 2, never wired into
   encode.py -- from a loose end to the highest-value work in the repo. It is
   worth a full step on the quality ladder (lam=25 -> lam=10, +0.7/+1.2 dB)
   because it allows sizing for the mean instead of the peak.

2. Pixel-exact is off the table at this bandwidth: lam=0 needs 92-97% of the
   pipe. The previous commit's "if SCSI sustains >=800 KB/s, ship transparent"
   conclusion only applies at roughly double the user's figure.

3. FINDINGS 5 said that because transfers are DMA, streaming "costs essentially
   no CPU" and the 68000 is "nearly idle". That is wrong. The HD63450 steals
   ~8 clocks per 16-bit word: 10-20% of the machine at the rates the profiles
   now use, on top of a 38% full-frame blit. Bandwidth and CPU are one budget.

Adds tools/encoder/profile_gen.py, which derives lam FROM a bandwidth figure
(accounting for audio, peak/mean and DMA steal) instead of reading it off the
knee of the RD curve, and docs/BENCHMARK.md covering how to actually measure
the subsystem -- including why MAME cannot answer the bandwidth question and
would be the same class of error as the FINDINGS 4 traps.

The 4 Mbps figure is user-supplied and its provenance is not recorded; every
profile now hangs off it.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 12:12:34 -07:00

6.1 KiB

Benchmarking the storage subsystem, and deriving profiles from it

Written session 2, in answer to "how do we benchmark the SCSI subsystem itself and base our performance profiles around that?"

The short answer

You cannot set a bitrate profile from MAME. MAME's x68k_hdc (SASI) and mb89352/cz6bs1 (SCSI) are functional models — they move the right bytes and raise the right interrupts, but they are not transfer-timing accurate. A throughput number out of MAME measures how fast the emulator's device model hands over a buffer, which is an artefact of MAME's scheduling, not of a Fujitsu MB89352 on a 10MHz bus.

So split the question in two, because they need different instruments:

question instrument what it settles
does our read path work at all? MAME correctness, IOCS vs direct SPC, DMA setup
what rate does the hardware sustain? derivation + real hardware the profile bitrates

Using MAME for the second is the same class of error as FINDINGS 4: a number that looks like a measurement but is an artefact of the apparatus.

Tier 1 — MAME: validate the path, not the speed

This is what tools/bench/ already does, and what is currently blocked (IOCS _B_READ returns -1 uniformly). Its value is that it proves the request/DMA/completion loop is correct before any of it is burned into 68000 player code.

Next moves, in order — the SCSI path was never tried and is more relevant to the target anyway:

  1. SCSI instead of SASI. -exp1 cz6bs1 -hard disk.chd, with exp1:cz6bs1:scsi:0 harddisk. Use IOCS _S_READ ($F5) rather than _B_READ ($46).
  2. Move the stack. SP=$8000 may sit on top of the IOCS work area in low RAM; put it at $200000+ (hypothesis 3 from session 1).
  3. Format the image. Hypothesis 1 — a raw image has no X68000 partition structure, so the IPL's boot scan never registers a drive and IOCS refuses. Needs a Human68k image, which this machine does not have.
  4. Bypass IOCS entirely and drive the MB89352 SPC registers directly. This is what the shipping player will do anyway, since we want DMA straight into a ring buffer with no OS in the path. If direct SPC works while IOCS does not, that is a complete answer to the blocker and we simply skip IOCS.

Record from MAME: bytes transferred, completion status, and whether DMA or PIO was used. Do not record KB/s and treat it as a hardware figure.

Tier 2 — derivation: the defensible ceiling

Already partly in FINDINGS 5. Bounds worth tightening from datasheets:

  • 68000 bus cycle: 4 clocks @ 10MHz, 16-bit => 5 MB/s absolute ceiling
  • HD63450 single-address DMA, ~8 clocks/word => ~2.5 MB/s practical ceiling
  • SCSI-1 asynchronous REQ/ACK handshake per byte, plus MB89352 FIFO depth => the real limiter, and the number we do not have from a primary source

The user's working figure is 4 Mbps = 488 KB/s, which sits sensibly between the derived DMA ceiling and observed period-drive rates. Provenance not yet recorded — worth pinning down, because every profile now hangs off it.

The coupling nobody had counted

Cycle-stealing DMA is not free DMA. At ~8 clocks per 16-bit word:

stream CPU stolen + full-frame blit (38.3%)
110 KB/s 4.5% 42.8%
250 KB/s 10.2% 48.5%
450 KB/s 18.4% 56.7%
488 KB/s 20.0% 58.3%

FINDINGS 5 concluded that because transfers are DMA, "streaming costs essentially no CPU". That is wrong. It costs up to a fifth of the machine at the rates we now care about. Bandwidth and CPU are one budget, not two.

Tier 3 — real hardware: the only thing that settles it

An X68000 (ACE/EXPERT for SASI, Super/XVI or a CZ-6BS1-equipped 10MHz machine for SCSI) with a BlueSCSI or SCSI2SD, which is the realistic deployment anyway and removes mechanical seek from the measurement.

The benchmark must measure what the player actually does, not a synthetic bulk read:

  1. Sequential read into a ring buffer, in the chunk size the player will use.
  2. With the decoder running — so DMA/CPU contention is included. An idle-CPU bulk read will overstate the sustained rate by roughly the blit percentage.
  3. Timed with the machine's own timer (MFP timer-C or the 1/100s system clock), not a stopwatch.
  4. Reported as sustained KB/s over >=30s, plus the worst 1-second window. The worst window is what the profile must survive, since a frame that arrives late is a dropped frame.

Deliverable: a .x executable and its source in tools/bench/, runnable on real hardware and reporting a single number.

Feeding the result back into the profiles

tools/encoder/profile_gen.py inverts the dependency — give it a bandwidth and it returns the lam that fits, from the MEASURED rate-distortion points in FINDINGS 17.4:

python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi

It accounts for the three things that eat the pipe before video sees any of it: audio (7.8 KB/s), peak-to-mean burstiness (measured 1.4-1.9x), and it reports the DMA cycle-steal so the CPU coupling stays visible.

At 4 Mbps it currently returns:

lam mean peak 00020 00146 CPU
today (no rate control) 25 194 KB/s 368 KB/s -1.22 dB -4.21 dB 53%
with rate control wired 10 305 KB/s 305 KB/s -0.52 dB -2.98 dB 51%

Rate control is worth a full step on the quality ladder — it is not a tidiness feature, it is the difference between sizing for the peak and sizing for the mean. That is the strongest argument yet for wiring up ratectl.py.

What would change the design

  • If sustained is much below 4 Mbps (say 2 Mbps / 244 KB/s): scsi collapses toward today's sasi, and the two profiles stop being meaningfully different. At that point reconsider 10 fps, or a narrower active area.
  • If sustained is much above (>=8 Mbps / 976 KB/s): lam=0 fits with margin and the port ships pixel-exact video on SCSI.
  • If DMA cannot be used and transfers fall back to PIO, the CPU cost rises from ~15% to something far larger and CPU becomes the binding constraint. This is the single worst outcome and is worth checking early in Tier 1.