Find the sustained action sequence: it breaks both profiles
The open risk since session 2 was "a sustained action sequence could still break the bitrate", with every clip measured so far being 1.2-1.7 s. Closed by measurement rather than by sampling clips by hand. 07_motion_survey.py scans a whole stream at 96x72 for the hottest sliding window of inter-frame difference. On 00223 the spread between the quietest and hottest sustained 10 s windows is 10.6x, which is the argument for not eyeballing it. Hottest is t=539.4s, the Singe endgame. There, with the fixed lam the CLI uses, sasi overshoots 110 -> 129.6 KB/s (+18%) and scsi 280 -> 373.8 KB/s (+34%). Rate control moves from "insurance, not a fix" to required, and is promoted above the full-disc survey. The bus is not broken -- 381.6 KB/s still fits the 488 KB/s figure -- so FINDINGS 21 survives, at 78% of the pipe instead of a comfortable margin. Three further corrections fall out: - The two largest streams on the disc are bonus material. 00216 is the feature with a burned-in commentary PiP; 00215 is the commentary. 00223 is the clean 9.4 min. A size-ranked survey would have encoded live action. - On hard content the 256-colour scene palette (31.33 dB) binds well before the X68000 display (40.81 dB); scsi is already within 0.51 dB of it. - FINDINGS 24.5's architecture question resolves to "both paths, chosen per frame": 30-53% of frames sit above the 70% crossover. Picking per frame costs a median 37.0% of the frame budget and caps at 53.6%. Reporting for this is wired into encode.py, which previously only printed a mean over all frames -- the one statistic that cannot answer a per-frame question. extract.py takes optional start/dur; 08_mode_map.py renders source | decoded | block-mode map to .webm. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -890,3 +890,116 @@ V1's output was snapshotted and passes `verify_frame256.py` unchanged: `256x512
|
||||
native, double-scan exact, active 256x192 pixel-exact, letterbox true black`,
|
||||
40.81 dB. So 68000 code drives the mode of FINDINGS 23 correctly, and 23.5 is
|
||||
now closed.
|
||||
|
||||
---
|
||||
|
||||
## 25. The sustained action sequence, found and measured (session 5)
|
||||
|
||||
STATUS has carried "a *sustained* action sequence is the one thing that could
|
||||
still break the bitrate" as the open risk since session 2. Every clip measured
|
||||
before this was 1.2-1.7 s. This section closes it: **it does break the profiles,
|
||||
though not the bus.**
|
||||
|
||||
### 25.1 The two largest streams on the disc are not game footage
|
||||
A survey that sorts 224 streams by size and encodes the biggest would have
|
||||
measured **live action**:
|
||||
|
||||
| stream | size | what it actually is |
|
||||
|---|---:|---|
|
||||
| 00216 | 3777 MB | the feature with a **burned-in picture-in-picture commentary** |
|
||||
| 00215 | 3475 MB | the commentary itself, full-screen live action |
|
||||
| **00223** | **1802 MB** | **clean animation, 9.4 min — the one to use** |
|
||||
|
||||
The PiP in 00216 is burned into video stream 0, not a selectable secondary
|
||||
stream, so there is no ffmpeg flag that recovers a clean frame from it. This
|
||||
extends FINDINGS 13's menu-vs-content warning: the classification needed is
|
||||
**content / menu / bonus**, and bonus material is the one that looks most like
|
||||
content by every cheap metric (size, duration, bitrate).
|
||||
|
||||
### 25.2 Picking the worst window by measurement, not by eye
|
||||
`tools/analysis/07_motion_survey.py` scans a whole stream at 96x72 and reports
|
||||
the highest-mean sliding window of inter-frame absolute difference. On 00223:
|
||||
|
||||
```
|
||||
6793 frames @12fps = 566.1s
|
||||
motion energy mean 9.40 median 5.60 p90 21.70 max 112.39
|
||||
hottest sustained 10s window: t = 539.4s (2.01x stream mean)
|
||||
quietest 10s window: t = 144.2s (0.19x stream mean)
|
||||
```
|
||||
|
||||
The 10.6x spread between the quietest and hottest sustained windows is the whole
|
||||
argument for not sampling clips by hand. `t = 539.4s` is the Singe endgame.
|
||||
|
||||
### 25.3 Both profiles overshoot on that window — rate control is now required
|
||||
Encoding those 120 frames at the shipping profiles, with the fixed `lam` the CLI
|
||||
currently uses:
|
||||
|
||||
| profile | target | measured | overshoot | PSNR | palette ceiling |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| `sasi` | 110 KB/s | **129.6 KB/s** | **+18%** | 27.82 dB | 31.33 dB |
|
||||
| `scsi` | 280 KB/s | **373.8 KB/s** | **+34%** | 30.81 dB | 31.33 dB |
|
||||
| *(00020 baseline, `sasi`)* | 110 KB/s | 108.0 KB/s | -2% | 36.94 dB | 39.90 dB |
|
||||
|
||||
**This reclassifies rate control from insurance to a requirement.** STATUS has
|
||||
had "wire rate control into `encode.py`" at priority 3-4 since session 2 with the
|
||||
note "no longer a blocker (FINDINGS 21)". That was true of the clips measured
|
||||
then. It is not true of this one. `ratectl.encode_rate_controlled()` already
|
||||
exists and builds a per-frame lam ladder; it has simply never been hooked up.
|
||||
|
||||
Note what did **not** break: 373.8 + 7.8 = 381.6 KB/s is still under the 488 KB/s
|
||||
working figure, so FINDINGS 21's ring-buffer conclusion survives — but at 78% of
|
||||
the pipe sustained over ten seconds rather than the comfortable margin implied by
|
||||
1.7 s clips.
|
||||
|
||||
### 25.4 The palette ceiling is content-dependent, and on hard content it binds
|
||||
The 256-colour scene palette costs **31.33 dB** on this window against **39.90 dB**
|
||||
on 00020 — 8.6 dB worse. Fire, lava and smoke gradients are exactly what a
|
||||
256-entry mediancut palette handles worst.
|
||||
|
||||
This inverts an assumption the project has been carrying. FINDINGS 23.3 put the
|
||||
X68000 display ceiling at 40.81 dB and treated it as comfortably clear of the
|
||||
codec's own error. On this content the **scene palette (31.33 dB), not the
|
||||
display hardware (40.81 dB), is the binding constraint** — and `scsi` is already
|
||||
within 0.51 dB of it. Spending bits to close that last half-dB is spending them
|
||||
against a ceiling that is not the display's.
|
||||
|
||||
### 25.5 `scsi` collapses to RAW under stress
|
||||
Mode distribution on this window is qualitatively different from anything
|
||||
measured before:
|
||||
|
||||
| profile | SKIP | V1 | V4 | RAW |
|
||||
|---|---:|---:|---:|---:|
|
||||
| `sasi` (lam=60) | 45.6% | 16.3% | 24.2% | 13.9% |
|
||||
| `scsi` (lam=10) | 26.2% | 5.5% | 7.1% | **61.2%** |
|
||||
| *00020, `sasi`* | 46.9% | 24.1% | 17.8% | 11.2% |
|
||||
|
||||
At `lam=10` the rate-distortion decision finds literal pixels cheaper than any
|
||||
codeword for 61% of blocks — the codebooks are simply not describing this
|
||||
content. That is the mechanism behind the +34% overshoot in 25.3, and it is a
|
||||
rate-control problem, not a codec-structure problem: the RD decision is behaving
|
||||
correctly for the lam it was given.
|
||||
|
||||
### 25.6 The decoder needs BOTH display paths, chosen per frame
|
||||
Applying FINDINGS 24.5's crossover to the real per-frame distribution:
|
||||
|
||||
| | median non-SKIP | p90 | frames over the 70% crossover |
|
||||
|---|---:|---:|---:|
|
||||
| `sasi`, Singe window | 48.4% | 82.8% | 36 / 120 (30%) |
|
||||
| `scsi`, Singe window | 70.8% | 92.4% | 64 / 120 (53%) |
|
||||
| `sasi`, 00020 | 54.0% | 88.8% | 3 / 14 (21%) |
|
||||
|
||||
Neither path wins outright: **30-53% of frames want the flat blit and the rest
|
||||
want direct-to-GVRAM.** A player that implements both and picks per frame — the
|
||||
mode headers are parsed before any pixel is written, so the count is free — pays
|
||||
a median of **37.0%** of the frame budget and is capped at **53.6%**. A player
|
||||
that implements only direct-to-GVRAM pays up to 76.6% and would miss frames on
|
||||
the scene cuts.
|
||||
|
||||
So the answer to 24.5 is "both", and the selection is a one-line comparison
|
||||
against a block count the decoder already has in hand.
|
||||
|
||||
### 25.7 What this does not measure
|
||||
One 10 s window of one stream, at fixed lam, with `_paint` still a Python loop.
|
||||
The full-disc survey is still not done, and the numbers above are the *worst*
|
||||
window rather than a distribution over content. What has changed is that the
|
||||
worst case is now a measurement rather than a worry.
|
||||
|
||||
Reference in New Issue
Block a user