User clarified the bandwidth figure is 4 Mbps (488 KB/s), not 4 MB/s -- ~8x
tighter than the previous commit reasoned against. Two consequences, plus a
correction to session 1.
1. The scsi profile committed in f0f2f80 DOES NOT FIT. Its mean is a
comfortable 52% of the pipe but it PEAKS at 96.4% (470.8 KB/s on scene
00020), and a frame that arrives late is a dropped frame, not a slow one.
Peak/mean is 1.4-1.9x even on 1.2-1.7s clips. Sizing a real-time stream on
the mean was the error. Flagged in STATUS rather than silently retuned,
because the fix is rate control, not a lower lam.
This promotes ratectl.py -- written in session 2, never wired into
encode.py -- from a loose end to the highest-value work in the repo. It is
worth a full step on the quality ladder (lam=25 -> lam=10, +0.7/+1.2 dB)
because it allows sizing for the mean instead of the peak.
2. Pixel-exact is off the table at this bandwidth: lam=0 needs 92-97% of the
pipe. The previous commit's "if SCSI sustains >=800 KB/s, ship transparent"
conclusion only applies at roughly double the user's figure.
3. FINDINGS 5 said that because transfers are DMA, streaming "costs essentially
no CPU" and the 68000 is "nearly idle". That is wrong. The HD63450 steals
~8 clocks per 16-bit word: 10-20% of the machine at the rates the profiles
now use, on top of a 38% full-frame blit. Bandwidth and CPU are one budget.
Adds tools/encoder/profile_gen.py, which derives lam FROM a bandwidth figure
(accounting for audio, peak/mean and DMA steal) instead of reading it off the
knee of the RD curve, and docs/BENCHMARK.md covering how to actually measure
the subsystem -- including why MAME cannot answer the bandwidth question and
would be the same class of error as the FINDINGS 4 traps.
The 4 Mbps figure is user-supplied and its provenance is not recorded; every
profile now hangs off it.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
6.1 KiB
Benchmarking the storage subsystem, and deriving profiles from it
Written session 2, in answer to "how do we benchmark the SCSI subsystem itself and base our performance profiles around that?"
The short answer
You cannot set a bitrate profile from MAME. MAME's x68k_hdc (SASI) and
mb89352/cz6bs1 (SCSI) are functional models — they move the right bytes
and raise the right interrupts, but they are not transfer-timing accurate. A
throughput number out of MAME measures how fast the emulator's device model
hands over a buffer, which is an artefact of MAME's scheduling, not of a
Fujitsu MB89352 on a 10MHz bus.
So split the question in two, because they need different instruments:
| question | instrument | what it settles |
|---|---|---|
| does our read path work at all? | MAME | correctness, IOCS vs direct SPC, DMA setup |
| what rate does the hardware sustain? | derivation + real hardware | the profile bitrates |
Using MAME for the second is the same class of error as FINDINGS 4: a number that looks like a measurement but is an artefact of the apparatus.
Tier 1 — MAME: validate the path, not the speed
This is what tools/bench/ already does, and what is currently blocked
(IOCS _B_READ returns -1 uniformly). Its value is that it proves the
request/DMA/completion loop is correct before any of it is burned into 68000
player code.
Next moves, in order — the SCSI path was never tried and is more relevant to the target anyway:
- SCSI instead of SASI.
-exp1 cz6bs1 -hard disk.chd, withexp1:cz6bs1:scsi:0 harddisk. Use IOCS_S_READ($F5) rather than_B_READ($46). - Move the stack.
SP=$8000may sit on top of the IOCS work area in low RAM; put it at $200000+ (hypothesis 3 from session 1). - Format the image. Hypothesis 1 — a raw image has no X68000 partition structure, so the IPL's boot scan never registers a drive and IOCS refuses. Needs a Human68k image, which this machine does not have.
- Bypass IOCS entirely and drive the MB89352 SPC registers directly. This is what the shipping player will do anyway, since we want DMA straight into a ring buffer with no OS in the path. If direct SPC works while IOCS does not, that is a complete answer to the blocker and we simply skip IOCS.
Record from MAME: bytes transferred, completion status, and whether DMA or PIO was used. Do not record KB/s and treat it as a hardware figure.
Tier 2 — derivation: the defensible ceiling
Already partly in FINDINGS 5. Bounds worth tightening from datasheets:
- 68000 bus cycle: 4 clocks @ 10MHz, 16-bit => 5 MB/s absolute ceiling
- HD63450 single-address DMA, ~8 clocks/word => ~2.5 MB/s practical ceiling
- SCSI-1 asynchronous REQ/ACK handshake per byte, plus MB89352 FIFO depth => the real limiter, and the number we do not have from a primary source
The user's working figure is 4 Mbps = 488 KB/s, which sits sensibly between the derived DMA ceiling and observed period-drive rates. Provenance not yet recorded — worth pinning down, because every profile now hangs off it.
The coupling nobody had counted
Cycle-stealing DMA is not free DMA. At ~8 clocks per 16-bit word:
| stream | CPU stolen | + full-frame blit (38.3%) |
|---|---|---|
| 110 KB/s | 4.5% | 42.8% |
| 250 KB/s | 10.2% | 48.5% |
| 450 KB/s | 18.4% | 56.7% |
| 488 KB/s | 20.0% | 58.3% |
FINDINGS 5 concluded that because transfers are DMA, "streaming costs essentially no CPU". That is wrong. It costs up to a fifth of the machine at the rates we now care about. Bandwidth and CPU are one budget, not two.
Tier 3 — real hardware: the only thing that settles it
An X68000 (ACE/EXPERT for SASI, Super/XVI or a CZ-6BS1-equipped 10MHz machine for SCSI) with a BlueSCSI or SCSI2SD, which is the realistic deployment anyway and removes mechanical seek from the measurement.
The benchmark must measure what the player actually does, not a synthetic bulk read:
- Sequential read into a ring buffer, in the chunk size the player will use.
- With the decoder running — so DMA/CPU contention is included. An idle-CPU bulk read will overstate the sustained rate by roughly the blit percentage.
- Timed with the machine's own timer (MFP timer-C or the 1/100s system clock), not a stopwatch.
- Reported as sustained KB/s over >=30s, plus the worst 1-second window. The worst window is what the profile must survive, since a frame that arrives late is a dropped frame.
Deliverable: a .x executable and its source in tools/bench/, runnable on
real hardware and reporting a single number.
Feeding the result back into the profiles
tools/encoder/profile_gen.py inverts the dependency — give it a bandwidth and
it returns the lam that fits, from the MEASURED rate-distortion points in
FINDINGS 17.4:
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
It accounts for the three things that eat the pipe before video sees any of it: audio (7.8 KB/s), peak-to-mean burstiness (measured 1.4-1.9x), and it reports the DMA cycle-steal so the CPU coupling stays visible.
At 4 Mbps it currently returns:
| lam | mean | peak | 00020 | 00146 | CPU | |
|---|---|---|---|---|---|---|
| today (no rate control) | 25 | 194 KB/s | 368 KB/s | -1.22 dB | -4.21 dB | 53% |
| with rate control wired | 10 | 305 KB/s | 305 KB/s | -0.52 dB | -2.98 dB | 51% |
Rate control is worth a full step on the quality ladder — it is not a
tidiness feature, it is the difference between sizing for the peak and sizing
for the mean. That is the strongest argument yet for wiring up ratectl.py.
What would change the design
- If sustained is much below 4 Mbps (say 2 Mbps / 244 KB/s):
scsicollapses toward today'ssasi, and the two profiles stop being meaningfully different. At that point reconsider 10 fps, or a narrower active area. - If sustained is much above (>=8 Mbps / 976 KB/s):
lam=0fits with margin and the port ships pixel-exact video on SCSI. - If DMA cannot be used and transfers fall back to PIO, the CPU cost rises from ~15% to something far larger and CPU becomes the binding constraint. This is the single worst outcome and is worth checking early in Tier 1.