Find the sustained action sequence: it breaks both profiles

The open risk since session 2 was "a sustained action sequence could still
break the bitrate", with every clip measured so far being 1.2-1.7 s. Closed by
measurement rather than by sampling clips by hand.

07_motion_survey.py scans a whole stream at 96x72 for the hottest sliding
window of inter-frame difference. On 00223 the spread between the quietest and
hottest sustained 10 s windows is 10.6x, which is the argument for not eyeballing
it. Hottest is t=539.4s, the Singe endgame.

There, with the fixed lam the CLI uses, sasi overshoots 110 -> 129.6 KB/s (+18%)
and scsi 280 -> 373.8 KB/s (+34%). Rate control moves from "insurance, not a
fix" to required, and is promoted above the full-disc survey. The bus is not
broken -- 381.6 KB/s still fits the 488 KB/s figure -- so FINDINGS 21 survives,
at 78% of the pipe instead of a comfortable margin.

Three further corrections fall out:

- The two largest streams on the disc are bonus material. 00216 is the feature
  with a burned-in commentary PiP; 00215 is the commentary. 00223 is the clean
  9.4 min. A size-ranked survey would have encoded live action.
- On hard content the 256-colour scene palette (31.33 dB) binds well before the
  X68000 display (40.81 dB); scsi is already within 0.51 dB of it.
- FINDINGS 24.5's architecture question resolves to "both paths, chosen per
  frame": 30-53% of frames sit above the 70% crossover. Picking per frame costs
  a median 37.0% of the frame budget and caps at 53.6%. Reporting for this is
  wired into encode.py, which previously only printed a mean over all frames --
  the one statistic that cannot answer a per-frame question.

extract.py takes optional start/dur; 08_mode_map.py renders source | decoded |
block-mode map to .webm.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
prosolis
2026-08-23 14:00:12 -07:00
parent 09a5a50065
commit e00264a058
7 changed files with 385 additions and 12 deletions
+67 -11
View File
@@ -96,7 +96,29 @@ rate-distortion curve, not two codecs.
3. **Reading the source frame is exactly half the blit cost** (V1 53.6% vs a
write-only floor V3 of 27.1%). That is what makes the architecture question
below live.
4. **The decoder architecture now hinges on one unmeasured number.** Writing
4. **That number is now measured, and the answer is "implement both paths".**
On the worst sustained window found on the disc, 30% of frames (`sasi`) to
53% (`scsi`) sit above the 70% crossover and want the flat blit; the rest
want direct-to-GVRAM. A player that picks per frame — the mode headers are
parsed before any pixel is written, so the count is free — pays a **median
37.0%** and is **capped at 53.6%**. FINDINGS 25.6.
5. **The sustained action sequence exists, was found by measurement, and breaks
both profiles.** `tools/analysis/07_motion_survey.py` scans a whole stream
for the hottest sliding window; on 00223 it is t=539.4s, the Singe endgame,
at 2.01x the stream mean. There, fixed-lam `sasi` overshoots 110 -> 129.6
KB/s (+18%) and `scsi` 280 -> 373.8 KB/s (+34%). **Rate control is no longer
insurance — it is required.** FINDINGS 25.3.
6. **The two largest streams on the disc are bonus material, not game footage.**
00216 is the feature with a burned-in commentary PiP; 00215 is the commentary
itself. **00223 (9.4 min) is the clean one.** A size-ranked survey would have
encoded live action. FINDINGS 25.1.
7. **On hard content the scene palette, not the display, is the binding
ceiling** — 31.33 dB on the Singe window against 39.90 dB on 00020 and 40.81
dB for the X68000 display. `scsi` is already within 0.51 dB of it.
FINDINGS 25.4.
### Superseded within session 5
4a. **The decoder architecture hinged on one unmeasured number.** Writing
codewords straight into GVRAM costs 76.6% of the frame budget for a *full*
frame (V4 — the 1024-byte stride kills the `movem.l` burst), but scales with
the non-SKIP block fraction and needs **no RAM reference frame at all**,
@@ -284,7 +306,11 @@ SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
## Next steps, in priority order
1. **Measure the non-SKIP block fraction.** *(new top priority, session 5)*
1. ~~**Measure the non-SKIP block fraction.**~~ **DONE, session 5** — FINDINGS
25.6. Answer: implement **both** display paths and pick per frame; median
37.0% of the frame budget, capped at 53.6%. Reporting is wired into
`encode.py`. Original framing kept below because the reasoning still governs
the decoder's inner loop:
FINDINGS 24.5: compose-in-RAM-then-blit costs a flat 53.6% of the frame
budget; decode-direct-to-GVRAM costs 76.6% x (fraction of blocks that are not
SKIP) and needs no RAM reference frame. **They cross at 70%.** Which side of
@@ -297,6 +323,15 @@ SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
scene-cut frame is ~100% non-SKIP and a held frame near 0%, and the mean of
those two is a number describing no actual frame.
1b. **Wire rate control into `encode.py`. NOW REQUIRED (was priority 4).**
FINDINGS 25.3: on the worst sustained window both profiles overshoot their
targets with the fixed `lam` the CLI uses — `sasi` by 18%, `scsi` by 34%.
The note that used to sit here, "no longer a blocker (FINDINGS 21)", was true
of the 1.2-1.7 s clips measured at the time and is not true of this one.
`ratectl.encode_rate_controlled()` already builds a per-frame lam ladder; it
has never been hooked up. Do this before the full-disc survey or the survey
measures an encoder nobody will ship.
2. **68000 decoder skeleton**, with the inner loop chosen by (1). Parse `DLX1`,
expand codebooks, blit per block mode. The display path is verified *by 68000
code* now (FINDINGS 24) and the harness pattern is `tools/bench/blit.s` +
@@ -312,16 +347,18 @@ SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
rejection, which was argued as "54% LZ4 with no room beside a 38% blit". The
conclusion gets *stronger*, not weaker, but the arithmetic should be restated.
3. **Full-disc survey.** Only 4 clips of 1.2-1.7 s out of 224 streams have been
measured, and 00146 already runs 23% hotter than 00020. A *sustained* action
sequence is the one thing that could still break the bitrate. Classify menu
vs content first (FINDINGS 13) or the averages are diluted by static menus.
**Vectorise `_paint` before this run** — it is a Python per-block loop.
Pairs naturally with (1): the same run produces both numbers.
3. **Full-disc survey.** Now scoped by session 5 rather than open-ended: the
worst *sustained* window is measured (FINDINGS 25), so what remains is the
distribution over content, not the worst case.
- Classify **content / menu / bonus** — not just menu vs content. FINDINGS
25.1: the two largest streams are bonus material and look like content by
size, duration and bitrate alike.
- Run `tools/analysis/07_motion_survey.py` per stream first; it is cheap
(96x72 greyscale) and gives a hot-window shortlist so the expensive encode
only runs where it matters.
- **Vectorise `_paint` before this run** — it is a Python per-block loop.
- Do it **after** rate control (1b), or it measures an encoder nobody ships.
4. **Wire rate control into `encode.py`.** No longer a blocker (FINDINGS 21), but
it is what gives a deterministic ceiling over content not yet measured, which
was the original reason for choosing VQ. Insurance, not a fix. Pairs with (1).
5. **Confirm DMA vs PIO in MAME** (see the benchmark section above) — cheap, and
the only thing that could still move CPU into the binding position.
6. **Resolve the framing question** (FINDINGS 12: crop vs squash vs wide).
@@ -425,3 +462,22 @@ V1's output. To check that snapshot is still pixel-exact:
Not added to `check.sh`: `check.sh` asserts pixel-exactness, and asserting wall
timings there would make the green-light check sensitive to host load.
## Reproducing the sustained-action result (session 5)
```
python3 tools/analysis/07_motion_survey.py 00223 10 # -> hottest window t=539.4s
python3 tools/encoder/extract.py 00223 tmp/fr_singe 12 crop 539.4 10.0
python3 tools/encoder/encode.py tmp/fr_singe tmp/singe_sasi.dlx --profile sasi
python3 tools/encoder/encode.py tmp/fr_singe tmp/singe_scsi.dlx --profile scsi
python3 tools/analysis/08_mode_map.py tmp/fr_singe tmp/singe_modes.webm \
--profile sasi --scale 2
```
`extract.py` now takes optional `[start_s] [dur_s]` — needed because 00223 is
9.4 min and the windows that stress the codec are seconds long.
`08_mode_map.py` renders palettised source | decoded | block-mode map at 12fps.
Output format follows the extension; **prefer `.webm`** — GIF re-quantises to
256 colours, which is a poor fit for output whose subject is colour fidelity,
and runs larger. It uses `yuv444p` because the mode map is flat saturated colour
on a 4-pixel grid and chroma subsampling smears exactly those edges.