Files
Dragon-s-Lair-X68k/docs/BENCHMARK.md
T
prosolis 64cd1ffd72 Handoff: reconcile docs and tooling with the corrections made this session
Session 2 reversed several of its own conclusions. The docs are append-only, so
a reader could land on a superseded section and act on it. This pass makes the
repo internally consistent.

Defects found and fixed in STATUS.md:
- claimed "Hybrid VQ with k=1024: no" as the answer to the linework question,
  directly contradicting FINDINGS 14, which rejected k=1024. Both profiles are
  k=256.
- malformed profile table (six column separators, five columns).
- next-steps list had two items numbered 3 and listed the full-disc survey
  twice.
- the disk-benchmark section still read CRITICAL-PATH with "if SCSI sustains
  >=800 KB/s, ship pixel-exact". That was written while the bandwidth figure
  was misread as 4 MB/s. At 4 Mbps pixel-exact needs 92-97% of the pipe and is
  not available, and the ring-buffer result means the design no longer hangs on
  the benchmark at all. Rewritten with what it IS still worth doing: confirming
  the 4 Mbps provenance, and confirming DMA is used rather than PIO.

FINDINGS now carries supersession blockquotes on 5, 8, 11, 17 and 18 pointing
at the sections that correct them. 18 is the dangerous one -- its peak-vs-
sustained test is reversed by 21 -- so it is marked DO NOT ACT ON THIS SECTION
while noting the per-frame data itself remains valid.

profile_gen.py had the same problem in code: it defaulted to the superseded
peak sizing and returned lam=25 where the docs say lam=10. The buffered test is
now the default and peak sizing is behind --size-for-peak as a bound only. A
tool that contradicts the findings is worse than no tool.

Also preserves the five measurement scripts that produced this session's
numbers as tools/analysis/05-09, following the session 1 precedent, and adds an
"explicitly abandoned -- do not re-propose" list to STATUS covering entropy
coding, k=1024 codebooks and flat 4x4 VQ.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-23 12:28:41 -07:00

142 lines
6.6 KiB
Markdown

# Benchmarking the storage subsystem, and deriving profiles from it
Written session 2, in answer to "how do we benchmark the SCSI subsystem itself
and base our performance profiles around that?"
## The short answer
**You cannot set a bitrate profile from MAME.** MAME's `x68k_hdc` (SASI) and
`mb89352`/`cz6bs1` (SCSI) are *functional* models — they move the right bytes
and raise the right interrupts, but they are not transfer-timing accurate. A
throughput number out of MAME measures how fast the emulator's device model
hands over a buffer, which is an artefact of MAME's scheduling, not of a
Fujitsu MB89352 on a 10MHz bus.
So split the question in two, because they need different instruments:
| question | instrument | what it settles |
|---|---|---|
| does our read path work at all? | MAME | correctness, IOCS vs direct SPC, DMA setup |
| what rate does the hardware sustain? | derivation + real hardware | the profile bitrates |
Using MAME for the second is the same class of error as FINDINGS 4: a number
that looks like a measurement but is an artefact of the apparatus.
## Tier 1 — MAME: validate the path, not the speed
This is what `tools/bench/` already does, and what is currently blocked
(`IOCS _B_READ` returns -1 uniformly). Its value is that it proves the
request/DMA/completion loop is correct before any of it is burned into 68000
player code.
Next moves, in order — the SCSI path was never tried and is more relevant to
the target anyway:
1. **SCSI instead of SASI.** `-exp1 cz6bs1 -hard disk.chd`, with
`exp1:cz6bs1:scsi:0 harddisk`. Use IOCS `_S_READ` ($F5) rather than
`_B_READ` ($46).
2. **Move the stack.** `SP=$8000` may sit on top of the IOCS work area in low
RAM; put it at $200000+ (hypothesis 3 from session 1).
3. **Format the image.** Hypothesis 1 — a raw image has no X68000 partition
structure, so the IPL's boot scan never registers a drive and IOCS refuses.
Needs a Human68k image, which this machine does not have.
4. **Bypass IOCS entirely** and drive the MB89352 SPC registers directly. This
is what the shipping player will do anyway, since we want DMA straight into
a ring buffer with no OS in the path. If direct SPC works while IOCS does
not, that is a complete answer to the blocker and we simply skip IOCS.
Record from MAME: bytes transferred, completion status, and whether DMA or PIO
was used. **Do not record KB/s and treat it as a hardware figure.**
## Tier 2 — derivation: the defensible ceiling
Already partly in FINDINGS 5. Bounds worth tightening from datasheets:
- 68000 bus cycle: 4 clocks @ 10MHz, 16-bit => **5 MB/s** absolute ceiling
- HD63450 single-address DMA, ~8 clocks/word => **~2.5 MB/s** practical ceiling
- SCSI-1 asynchronous REQ/ACK handshake per byte, plus MB89352 FIFO depth
=> the real limiter, and the number we do not have from a primary source
The user's working figure is **4 Mbps = 488 KB/s**, which sits sensibly between
the derived DMA ceiling and observed period-drive rates. **Provenance not yet
recorded — worth pinning down, because every profile now hangs off it.**
### The coupling nobody had counted
Cycle-stealing DMA is not free DMA. At ~8 clocks per 16-bit word:
| stream | CPU stolen | + full-frame blit (38.3%) |
|---|---|---|
| 110 KB/s | 4.5% | 42.8% |
| 250 KB/s | 10.2% | 48.5% |
| 450 KB/s | 18.4% | 56.7% |
| 488 KB/s | 20.0% | 58.3% |
FINDINGS 5 concluded that because transfers are DMA, "streaming costs
essentially no CPU". **That is wrong.** It costs up to a fifth of the machine at
the rates we now care about. Bandwidth and CPU are one budget, not two.
## Tier 3 — real hardware: the only thing that settles it
An X68000 (ACE/EXPERT for SASI, Super/XVI or a CZ-6BS1-equipped 10MHz machine
for SCSI) with a **BlueSCSI or SCSI2SD**, which is the realistic deployment
anyway and removes mechanical seek from the measurement.
The benchmark must measure **what the player actually does**, not a synthetic
bulk read:
1. Sequential read into a ring buffer, in the chunk size the player will use.
2. **With the decoder running** — so DMA/CPU contention is included. An idle-CPU
bulk read will overstate the sustained rate by roughly the blit percentage.
3. Timed with the machine's own timer (MFP timer-C or the 1/100s system clock),
not a stopwatch.
4. Reported as sustained KB/s over >=30s, plus the worst 1-second window. The
worst window is what the profile must survive, since a frame that arrives
late is a dropped frame.
Deliverable: a `.x` executable and its source in `tools/bench/`, runnable on
real hardware and reporting a single number.
## Feeding the result back into the profiles
`tools/encoder/profile_gen.py` inverts the dependency — give it a bandwidth and
it returns the lam that fits, from the MEASURED rate-distortion points in
FINDINGS 17.4:
```
python3 tools/encoder/profile_gen.py --bw-mbps 4 --name scsi
```
It accounts for what eats the pipe before video sees any of it: audio
(7.8 KB/s), the buffering condition, and it reports the DMA cycle-steal so the
CPU coupling stays visible.
At 4 Mbps it returns:
| sizing rule | lam | mean | 00020 | 00146 | CPU |
|---|---|---|---|---|---|
| **buffered (default, FINDINGS 21)** | **10** | 305 KB/s | -0.52 dB | -2.98 dB | 51% |
| `--size-for-peak` (FINDINGS 18, superseded) | 25 | 194 KB/s | -1.22 dB | -4.21 dB | 46% |
The default is the buffered test: cumulative demand vs cumulative supply.
Ring-buffer simulation gives **zero required prefill** for every measured scene,
so `lam=10` ships without rate control. `--size-for-peak` reproduces the earlier
pessimistic sizing and is kept only as a bound.
**Rate control is therefore insurance, not a fix.** Its value is a deterministic
ceiling over the 220 streams not yet measured — see the survey caveat below.
## What would change the design
- **If sustained is much below 4 Mbps** (say 2 Mbps / 244 KB/s): `scsi`
collapses toward today's `sasi`, and the two profiles stop being meaningfully
different. At that point reconsider 10 fps, or a narrower active area.
- **If sustained is much above** (>=8 Mbps / 976 KB/s): `lam=0` fits with
margin and the port ships **pixel-exact** video on SCSI. At the current
4 Mbps figure this is NOT available — `lam=0` needs 92-97% of the pipe.
- **If the full-disc survey finds a sustained action sequence hotter than
00146** (313 KB/s mean, the worst of 4 clips sampled): that is the scenario
rate control exists for, and the reason to wire it up before the survey run.
- **If DMA cannot be used** and transfers fall back to PIO, the CPU cost rises
from ~15% to something far larger and CPU becomes the binding constraint.
This is the single worst outcome and is worth checking early in Tier 1.