Find the sustained action sequence: it breaks both profiles

The open risk since session 2 was "a sustained action sequence could still
break the bitrate", with every clip measured so far being 1.2-1.7 s. Closed by
measurement rather than by sampling clips by hand.

07_motion_survey.py scans a whole stream at 96x72 for the hottest sliding
window of inter-frame difference. On 00223 the spread between the quietest and
hottest sustained 10 s windows is 10.6x, which is the argument for not eyeballing
it. Hottest is t=539.4s, the Singe endgame.

There, with the fixed lam the CLI uses, sasi overshoots 110 -> 129.6 KB/s (+18%)
and scsi 280 -> 373.8 KB/s (+34%). Rate control moves from "insurance, not a
fix" to required, and is promoted above the full-disc survey. The bus is not
broken -- 381.6 KB/s still fits the 488 KB/s figure -- so FINDINGS 21 survives,
at 78% of the pipe instead of a comfortable margin.

Three further corrections fall out:

- The two largest streams on the disc are bonus material. 00216 is the feature
  with a burned-in commentary PiP; 00215 is the commentary. 00223 is the clean
  9.4 min. A size-ranked survey would have encoded live action.
- On hard content the 256-colour scene palette (31.33 dB) binds well before the
  X68000 display (40.81 dB); scsi is already within 0.51 dB of it.
- FINDINGS 24.5's architecture question resolves to "both paths, chosen per
  frame": 30-53% of frames sit above the 70% crossover. Picking per frame costs
  a median 37.0% of the frame budget and caps at 53.6%. Reporting for this is
  wired into encode.py, which previously only printed a mean over all frames --
  the one statistic that cannot answer a per-frame question.

extract.py takes optional start/dur; 08_mode_map.py renders source | decoded |
block-mode map to .webm.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
prosolis
2026-08-23 14:00:12 -07:00
parent 09a5a50065
commit e00264a058
7 changed files with 385 additions and 12 deletions
+113
View File
@@ -890,3 +890,116 @@ V1's output was snapshotted and passes `verify_frame256.py` unchanged: `256x512
native, double-scan exact, active 256x192 pixel-exact, letterbox true black`,
40.81 dB. So 68000 code drives the mode of FINDINGS 23 correctly, and 23.5 is
now closed.
---
## 25. The sustained action sequence, found and measured (session 5)
STATUS has carried "a *sustained* action sequence is the one thing that could
still break the bitrate" as the open risk since session 2. Every clip measured
before this was 1.2-1.7 s. This section closes it: **it does break the profiles,
though not the bus.**
### 25.1 The two largest streams on the disc are not game footage
A survey that sorts 224 streams by size and encodes the biggest would have
measured **live action**:
| stream | size | what it actually is |
|---|---:|---|
| 00216 | 3777 MB | the feature with a **burned-in picture-in-picture commentary** |
| 00215 | 3475 MB | the commentary itself, full-screen live action |
| **00223** | **1802 MB** | **clean animation, 9.4 min — the one to use** |
The PiP in 00216 is burned into video stream 0, not a selectable secondary
stream, so there is no ffmpeg flag that recovers a clean frame from it. This
extends FINDINGS 13's menu-vs-content warning: the classification needed is
**content / menu / bonus**, and bonus material is the one that looks most like
content by every cheap metric (size, duration, bitrate).
### 25.2 Picking the worst window by measurement, not by eye
`tools/analysis/07_motion_survey.py` scans a whole stream at 96x72 and reports
the highest-mean sliding window of inter-frame absolute difference. On 00223:
```
6793 frames @12fps = 566.1s
motion energy mean 9.40 median 5.60 p90 21.70 max 112.39
hottest sustained 10s window: t = 539.4s (2.01x stream mean)
quietest 10s window: t = 144.2s (0.19x stream mean)
```
The 10.6x spread between the quietest and hottest sustained windows is the whole
argument for not sampling clips by hand. `t = 539.4s` is the Singe endgame.
### 25.3 Both profiles overshoot on that window — rate control is now required
Encoding those 120 frames at the shipping profiles, with the fixed `lam` the CLI
currently uses:
| profile | target | measured | overshoot | PSNR | palette ceiling |
|---|---:|---:|---:|---:|---:|
| `sasi` | 110 KB/s | **129.6 KB/s** | **+18%** | 27.82 dB | 31.33 dB |
| `scsi` | 280 KB/s | **373.8 KB/s** | **+34%** | 30.81 dB | 31.33 dB |
| *(00020 baseline, `sasi`)* | 110 KB/s | 108.0 KB/s | -2% | 36.94 dB | 39.90 dB |
**This reclassifies rate control from insurance to a requirement.** STATUS has
had "wire rate control into `encode.py`" at priority 3-4 since session 2 with the
note "no longer a blocker (FINDINGS 21)". That was true of the clips measured
then. It is not true of this one. `ratectl.encode_rate_controlled()` already
exists and builds a per-frame lam ladder; it has simply never been hooked up.
Note what did **not** break: 373.8 + 7.8 = 381.6 KB/s is still under the 488 KB/s
working figure, so FINDINGS 21's ring-buffer conclusion survives — but at 78% of
the pipe sustained over ten seconds rather than the comfortable margin implied by
1.7 s clips.
### 25.4 The palette ceiling is content-dependent, and on hard content it binds
The 256-colour scene palette costs **31.33 dB** on this window against **39.90 dB**
on 00020 — 8.6 dB worse. Fire, lava and smoke gradients are exactly what a
256-entry mediancut palette handles worst.
This inverts an assumption the project has been carrying. FINDINGS 23.3 put the
X68000 display ceiling at 40.81 dB and treated it as comfortably clear of the
codec's own error. On this content the **scene palette (31.33 dB), not the
display hardware (40.81 dB), is the binding constraint** — and `scsi` is already
within 0.51 dB of it. Spending bits to close that last half-dB is spending them
against a ceiling that is not the display's.
### 25.5 `scsi` collapses to RAW under stress
Mode distribution on this window is qualitatively different from anything
measured before:
| profile | SKIP | V1 | V4 | RAW |
|---|---:|---:|---:|---:|
| `sasi` (lam=60) | 45.6% | 16.3% | 24.2% | 13.9% |
| `scsi` (lam=10) | 26.2% | 5.5% | 7.1% | **61.2%** |
| *00020, `sasi`* | 46.9% | 24.1% | 17.8% | 11.2% |
At `lam=10` the rate-distortion decision finds literal pixels cheaper than any
codeword for 61% of blocks — the codebooks are simply not describing this
content. That is the mechanism behind the +34% overshoot in 25.3, and it is a
rate-control problem, not a codec-structure problem: the RD decision is behaving
correctly for the lam it was given.
### 25.6 The decoder needs BOTH display paths, chosen per frame
Applying FINDINGS 24.5's crossover to the real per-frame distribution:
| | median non-SKIP | p90 | frames over the 70% crossover |
|---|---:|---:|---:|
| `sasi`, Singe window | 48.4% | 82.8% | 36 / 120 (30%) |
| `scsi`, Singe window | 70.8% | 92.4% | 64 / 120 (53%) |
| `sasi`, 00020 | 54.0% | 88.8% | 3 / 14 (21%) |
Neither path wins outright: **30-53% of frames want the flat blit and the rest
want direct-to-GVRAM.** A player that implements both and picks per frame — the
mode headers are parsed before any pixel is written, so the count is free — pays
a median of **37.0%** of the frame budget and is capped at **53.6%**. A player
that implements only direct-to-GVRAM pays up to 76.6% and would miss frames on
the scene cuts.
So the answer to 24.5 is "both", and the selection is a one-line comparison
against a block count the decoder already has in hand.
### 25.7 What this does not measure
One 10 s window of one stream, at fixed lam, with `_paint` still a Python loop.
The full-disc survey is still not done, and the numbers above are the *worst*
window rather than a distribution over content. What has changed is that the
worst case is now a measurement rather than a worry.