Files
Dragon-s-Lair-X68k/docs/BENCHMARK.md
T
prosolis 64cd1ffd72 Handoff: reconcile docs and tooling with the corrections made this session
Session 2 reversed several of its own conclusions. The docs are append-only, so
a reader could land on a superseded section and act on it. This pass makes the
repo internally consistent.

Defects found and fixed in STATUS.md:
- claimed "Hybrid VQ with k=1024: no" as the answer to the linework question,
  directly contradicting FINDINGS 14, which rejected k=1024. Both profiles are
  k=256.
- malformed profile table (six column separators, five columns).
- next-steps list had two items numbered 3 and listed the full-disc survey
  twice.
- the disk-benchmark section still read CRITICAL-PATH with "if SCSI sustains
  >=800 KB/s, ship pixel-exact". That was written while the bandwidth figure
  was misread as 4 MB/s. At 4 Mbps pixel-exact needs 92-97% of the pipe and is
  not available, and the ring-buffer result means the design no longer hangs on
  the benchmark at all. Rewritten with what it IS still worth doing: confirming
  the 4 Mbps provenance, and confirming DMA is used rather than PIO.

FINDINGS now carries supersession blockquotes on 5, 8, 11, 17 and 18 pointing
at the sections that correct them. 18 is the dangerous one -- its peak-vs-
sustained test is reversed by 21 -- so it is marked DO NOT ACT ON THIS SECTION
while noting the per-frame data itself remains valid.

profile_gen.py had the same problem in code: it defaulted to the superseded
peak sizing and returned lam=25 where the docs say lam=10. The buffered test is
now the default and peak sizing is behind --size-for-peak as a bound only. A
tool that contradicts the findings is worse than no tool.

Also preserves the five measurement scripts that produced this session's
numbers as tools/analysis/05-09, following the session 1 precedent, and adds an
"explicitly abandoned -- do not re-propose" list to STATUS covering entropy
coding, k=1024 codebooks and flat 4x4 VQ.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 12:28:41 -07:00

6.6 KiB

Benchmarking the storage subsystem, and deriving profiles from it

Written session 2, in answer to "how do we benchmark the SCSI subsystem itself and base our performance profiles around that?"

The short answer

You cannot set a bitrate profile from MAME. MAME's x68k_hdc (SASI) and mb89352/cz6bs1 (SCSI) are functional models — they move the right bytes and raise the right interrupts, but they are not transfer-timing accurate. A throughput number out of MAME measures how fast the emulator's device model hands over a buffer, which is an artefact of MAME's scheduling, not of a Fujitsu MB89352 on a 10MHz bus.

So split the question in two, because they need different instruments:

question instrument what it settles
does our read path work at all? MAME correctness, IOCS vs direct SPC, DMA setup
what rate does the hardware sustain? derivation + real hardware the profile bitrates

Using MAME for the second is the same class of error as FINDINGS 4: a number that looks like a measurement but is an artefact of the apparatus.

Tier 1 — MAME: validate the path, not the speed

This is what tools/bench/ already does, and what is currently blocked (IOCS _B_READ returns -1 uniformly). Its value is that it proves the request/DMA/completion loop is correct before any of it is burned into 68000 player code.

Next moves, in order — the SCSI path was never tried and is more relevant to the target anyway:

  1. SCSI instead of SASI. -exp1 cz6bs1 -hard disk.chd, with exp1:cz6bs1:scsi:0 harddisk. Use IOCS _S_READ ($F5) rather than _B_READ ($46).
  2. Move the stack. SP=$8000 may sit on top of the IOCS work area in low RAM; put it at $200000+ (hypothesis 3 from session 1).
  3. Format the image. Hypothesis 1 — a raw image has no X68000 partition structure, so the IPL's boot scan never registers a drive and IOCS refuses. Needs a Human68k image, which this machine does not have.
  4. Bypass IOCS entirely and drive the MB89352 SPC registers directly. This is what the shipping player will do anyway, since we want DMA straight into a ring buffer with no OS in the path. If direct SPC works while IOCS does not, that is a complete answer to the blocker and we simply skip IOCS.

Record from MAME: bytes transferred, completion status, and whether DMA or PIO was used. Do not record KB/s and treat it as a hardware figure.

Tier 2 — derivation: the defensible ceiling

Already partly in FINDINGS 5. Bounds worth tightening from datasheets:

  • 68000 bus cycle: 4 clocks @ 10MHz, 16-bit => 5 MB/s absolute ceiling
  • HD63450 single-address DMA, ~8 clocks/word => ~2.5 MB/s practical ceiling
  • SCSI-1 asynchronous REQ/ACK handshake per byte, plus MB89352 FIFO depth => the real limiter, and the number we do not have from a primary source

The user's working figure is 4 Mbps = 488 KB/s, which sits sensibly between the derived DMA ceiling and observed period-drive rates. Provenance not yet recorded — worth pinning down, because every profile now hangs off it.

The coupling nobody had counted

Cycle-stealing DMA is not free DMA. At ~8 clocks per 16-bit word:

stream CPU stolen + full-frame blit (38.3%)
110 KB/s 4.5% 42.8%
250 KB/s 10.2% 48.5%
450 KB/s 18.4% 56.7%
488 KB/s 20.0% 58.3%

FINDINGS 5 concluded that because transfers are DMA, "streaming costs essentially no CPU". That is wrong. It costs up to a fifth of the machine at the rates we now care about. Bandwidth and CPU are one budget, not two.

Tier 3 — real hardware: the only thing that settles it

An X68000 (ACE/EXPERT for SASI, Super/XVI or a CZ-6BS1-equipped 10MHz machine for SCSI) with a BlueSCSI or SCSI2SD, which is the realistic deployment anyway and removes mechanical seek from the measurement.

The benchmark must measure what the player actually does, not a synthetic bulk read:

  1. Sequential read into a ring buffer, in the chunk size the player will use.
  2. With the decoder running — so DMA/CPU contention is included. An idle-CPU bulk read will overstate the sustained rate by roughly the blit percentage.
  3. Timed with the machine's own timer (MFP timer-C or the 1/100s system clock), not a stopwatch.
  4. Reported as sustained KB/s over >=30s, plus the worst 1-second window. The worst window is what the profile must survive, since a frame that arrives late is a dropped frame.

Deliverable: a .x executable and its source in tools/bench/, runnable on real hardware and reporting a single number.

Feeding the result back into the profiles

tools/encoder/profile_gen.py inverts the dependency — give it a bandwidth and it returns the lam that fits, from the MEASURED rate-distortion points in FINDINGS 17.4:

python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi

It accounts for what eats the pipe before video sees any of it: audio (7.8 KB/s), the buffering condition, and it reports the DMA cycle-steal so the CPU coupling stays visible.

At 4 Mbps it returns:

sizing rule lam mean 00020 00146 CPU
buffered (default, FINDINGS 21) 10 305 KB/s -0.52 dB -2.98 dB 51%
--size-for-peak (FINDINGS 18, superseded) 25 194 KB/s -1.22 dB -4.21 dB 46%

The default is the buffered test: cumulative demand vs cumulative supply. Ring-buffer simulation gives zero required prefill for every measured scene, so lam=10 ships without rate control. --size-for-peak reproduces the earlier pessimistic sizing and is kept only as a bound.

Rate control is therefore insurance, not a fix. Its value is a deterministic ceiling over the 220 streams not yet measured — see the survey caveat below.

What would change the design

  • If sustained is much below 4 Mbps (say 2 Mbps / 244 KB/s): scsi collapses toward today's sasi, and the two profiles stop being meaningfully different. At that point reconsider 10 fps, or a narrower active area.
  • If sustained is much above (>=8 Mbps / 976 KB/s): lam=0 fits with margin and the port ships pixel-exact video on SCSI. At the current 4 Mbps figure this is NOT available — lam=0 needs 92-97% of the pipe.
  • If the full-disc survey finds a sustained action sequence hotter than 00146 (313 KB/s mean, the worst of 4 clips sampled): that is the scenario rate control exists for, and the reason to wire it up before the survey run.
  • If DMA cannot be used and transfers fall back to PIO, the CPU cost rises from ~15% to something far larger and CPU becomes the binding constraint. This is the single worst outcome and is worth checking early in Tier 1.