# Benchmarking the storage subsystem, and deriving profiles from it Written session 2, in answer to "how do we benchmark the SCSI subsystem itself and base our performance profiles around that?" ## The short answer **You cannot set a bitrate profile from MAME.** MAME's `x68k_hdc` (SASI) and `mb89352`/`cz6bs1` (SCSI) are *functional* models — they move the right bytes and raise the right interrupts, but they are not transfer-timing accurate. A throughput number out of MAME measures how fast the emulator's device model hands over a buffer, which is an artefact of MAME's scheduling, not of a Fujitsu MB89352 on a 10MHz bus. So split the question in two, because they need different instruments: | question | instrument | what it settles | |---|---|---| | does our read path work at all? | MAME | correctness, IOCS vs direct SPC, DMA setup | | what rate does the hardware sustain? | derivation + real hardware | the profile bitrates | Using MAME for the second is the same class of error as FINDINGS 4: a number that looks like a measurement but is an artefact of the apparatus. ## Tier 1 — MAME: validate the path, not the speed This is what `tools/bench/` already does, and what is currently blocked (`IOCS _B_READ` returns -1 uniformly). Its value is that it proves the request/DMA/completion loop is correct before any of it is burned into 68000 player code. Next moves, in order — the SCSI path was never tried and is more relevant to the target anyway: 1. **SCSI instead of SASI.** `-exp1 cz6bs1 -hard disk.chd`, with `exp1:cz6bs1:scsi:0 harddisk`. Use IOCS `_S_READ` ($F5) rather than `_B_READ` ($46). 2. **Move the stack.** `SP=$8000` may sit on top of the IOCS work area in low RAM; put it at $200000+ (hypothesis 3 from session 1). 3. **Format the image.** Hypothesis 1 — a raw image has no X68000 partition structure, so the IPL's boot scan never registers a drive and IOCS refuses. Needs a Human68k image, which this machine does not have. 4. **Bypass IOCS entirely** and drive the MB89352 SPC registers directly. This is what the shipping player will do anyway, since we want DMA straight into a ring buffer with no OS in the path. If direct SPC works while IOCS does not, that is a complete answer to the blocker and we simply skip IOCS. Record from MAME: bytes transferred, completion status, and whether DMA or PIO was used. **Do not record KB/s and treat it as a hardware figure.** ## Tier 2 — derivation: the defensible ceiling Already partly in FINDINGS 5. Bounds worth tightening from datasheets: - 68000 bus cycle: 4 clocks @ 10MHz, 16-bit => **5 MB/s** absolute ceiling - HD63450 single-address DMA, **5 clocks/BYTE** => **2.0 MB/s** practical ceiling (dual-address is 9 clocks/byte => 1.11 MB/s). CORRECTED session 14: this line read "~8 clocks/word => ~2.5 MB/s", which charged a byte-wide SPC per word. FINDINGS 43. - SCSI-1 asynchronous REQ/ACK handshake per byte, plus MB89352 FIFO depth => the real limiter, and the number we do not have from a primary source **RETIRED, session 18 (USER DECISION).** This document used to name a working figure of "4 Mbps" here and note that every profile hung off it. It was never a bus measurement — user-supplied, no provenance, and 10% of SCSI-1's asynchronous rating (FINDINGS 42.1). It has been removed as a default from every analysis tool and from `tools/bench/stream.lua`; the tools now REQUIRE an explicit rate, so nothing can be scored against a figure the scorer never restates. **There is no working delivery figure. That is the honest state, and it is the point:** the rate is a property of the medium, the medium is a BlueSCSI, and it has not been measured. `tools/analysis/19_ring_stream.py` reports the **zero-prefill pipe** — the rate a medium must clear for a given container to need no prefill at all — which is the threshold a measurement should be taken against. For the session-14 candidate that is **513.2 KB/s** (FINDINGS 49.5). One place still carries the old number: `GATE_SPAN_KBPS` in `tools/bench/check.sh`, because the gate container was *encoded* with it and every per-block and span constant in FINDINGS 41/43/45/49 is fitted to that container. It is a container recipe, not a claim about any medium. ### The coupling nobody had counted Cycle-stealing DMA is not free DMA. At ~8 clocks per 16-bit word: | stream | CPU stolen | + full-frame blit (38.3%) | |---|---|---| | 110 KB/s | 4.5% | 42.8% | | 250 KB/s | 10.2% | 48.5% | | 450 KB/s | 18.4% | 56.7% | | ~490 KB/s | 20.0% | 58.3% | FINDINGS 5 concluded that because transfers are DMA, "streaming costs essentially no CPU". **That is wrong.** It costs up to a fifth of the machine at the rates we now care about. Bandwidth and CPU are one budget, not two. ## Tier 3 — real hardware: the only thing that settles it An X68000 (ACE/EXPERT for SASI, Super/XVI or a CZ-6BS1-equipped 10MHz machine for SCSI) with a **BlueSCSI or SCSI2SD**, which is the realistic deployment anyway and removes mechanical seek from the measurement. The benchmark must measure **what the player actually does**, not a synthetic bulk read: 1. Sequential read into a ring buffer, in the chunk size the player will use. 2. **With the decoder running** — so DMA/CPU contention is included. An idle-CPU bulk read will overstate the sustained rate by roughly the blit percentage. 3. Timed with the machine's own timer (MFP timer-C or the 1/100s system clock), not a stopwatch. 4. Reported as sustained KB/s over >=30s, plus the worst 1-second window. The worst window is what the profile must survive, since a frame that arrives late is a dropped frame. Deliverable: a `.x` executable and its source in `tools/bench/`, runnable on real hardware and reporting a single number. ## Feeding the result back into the profiles `tools/encoder/profile_gen.py` inverts the dependency — give it a bandwidth and it returns the lam that fits, from the MEASURED rate-distortion points in FINDINGS 17.4: ``` python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi ``` It accounts for what eats the pipe before video sees any of it: audio (7.8 KB/s), the buffering condition, and it reports the DMA cycle-steal so the CPU coupling stays visible. At 4 Mbps it returns: | sizing rule | lam | mean | 00020 | 00146 | CPU | |---|---|---|---|---|---| | **buffered (default, FINDINGS 21)** | **10** | 305 KB/s | -0.52 dB | -2.98 dB | 51% | | `--size-for-peak` (FINDINGS 18, superseded) | 25 | 194 KB/s | -1.22 dB | -4.21 dB | 46% | The default is the buffered test: cumulative demand vs cumulative supply. Ring-buffer simulation gives **zero required prefill** for every measured scene, so `lam=10` ships without rate control. `--size-for-peak` reproduces the earlier pessimistic sizing and is kept only as a bound. **Rate control is therefore insurance, not a fix.** Its value is a deterministic ceiling over the 220 streams not yet measured — see the survey caveat below. ## What would change the design - **If sustained is much below 4 Mbps** (say 2 Mbps / 244 KB/s): `scsi` collapses toward today's `sasi`, and the two profiles stop being meaningfully different. At that point reconsider 10 fps, or a narrower active area. - **If sustained is much above** (>=8 Mbps / 976 KB/s): `lam=0` fits with margin and the port ships **pixel-exact** video on SCSI. At the current 4 Mbps figure this is NOT available — `lam=0` needs 92-97% of the pipe. - **If the full-disc survey finds a sustained action sequence hotter than 00146** (313 KB/s mean, the worst of 4 clips sampled): that is the scenario rate control exists for, and the reason to wire it up before the survey run. - **If DMA cannot be used** and transfers fall back to PIO, the CPU cost rises from ~15% to something far larger and CPU becomes the binding constraint. This is the single worst outcome and is worth checking early in Tier 1.