Price cycles in the mode decision: 37 misses become 1, for 0.26 dB
The decoder has been CPU-bound since FINDINGS 28 while the mode decision
minimised D + lam*R -- distortion against BYTES. decide() now minimises
D + lam*bytes + mu*cycles, and ratectl bisects mu per frame against the
833,333-cycle budget with the lam bisection nested inside it. On the worst
sustained window:
sasi 27.22 -> 26.95 dB, 109.5 -> 109.4 KB/s, 37/120 misses -> 1
scsi 29.90 -> 29.27 dB, 280.0 -> 278.6 KB/s, 51/120 misses -> 1
Bitrate does not move: the byte controller still binds, and mu changes WHICH
modes are bought. V4 is what it stops buying -- 25.2 -> 20.3% of blocks at sasi
and 15.0 -> 5.3% at scsi, where RAW takes it. That is 28.8's inversion in
practice: RAW is dearer in bytes and cheaper in cycles, so only the byte-rich
profile can buy its way out of V4.
Three things worth knowing beyond the headline:
- The one frame that still misses, at both profiles, is FRAME 0 -- no previous
reconstruction, so 100% changed by definition, which is also what a scene
cut is. It comes out at the all-V1 floor of 110.6% and is emitted late on
purpose. Freezing a cut to make a deadline is the worse failure.
- 28.7's "11 frames are impossible" was too pessimistic. That floor held the
SKIP set fixed and asked how cheaply the drawn blocks could be drawn; the
real decision can also MOVE a block to SKIP, which above ~90% non-SKIP is
the only lever left.
- SKIP's price depends on its neighbours (13.25 cycles clustered, 45 mixed),
which a per-block lagrangian cannot see. The way out is that the two uses
need not share a cost function: a ranking constant inside decide(), the
exact clustered rule for the frame-level bisection. vq_hybrid.cycles() is
now the one definition of that rule and 11_cpu_budget.py imports it.
Gated: 09_ratectl_drift.py runs both controllers, both 0/120 drifting frames.
The cost-aware container decodes pixel-exact on the 68000 (120 frames). ON by
default in encode.py; --no-cpu-fit restores session 7. check.sh ALL GREEN.
Still a model, not a measurement, for THIS container: FINDINGS 31's cycle
figures come from vq_hybrid.cycles (within 1 point of the 68000 on four frames
of the session-7 container). Timing this one on the machine is step 1 of the
next session -- it was started and killed for time, and it is slow.
FINDINGS 31. tools/analysis/13_cpu_ratectl.py.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -39,7 +39,9 @@ tools/analysis/ measurement scripts, numbered in the order they were written
|
||||
11 scores a container against the MEASURED per-mode block
|
||||
costs without needing MAME; 12 prices the literal-span mode of
|
||||
FINDINGS 30 against those same mode maps, and prints whether a
|
||||
scene cut still fits at 12fps.
|
||||
scene cut still fits at 12fps; 13 measures what fitting the
|
||||
CPU budget costs in dB (FINDINGS 31) and caches H.build so the
|
||||
search loop is seconds, not minutes.
|
||||
tools/bench/ MAME Lua injection harness + 68000 benchmark sources.
|
||||
`check.sh` re-runs both display regression tests (~40 s).
|
||||
`blit.s`/`blit.lua` time the full-frame GVRAM blit on the
|
||||
@@ -70,7 +72,14 @@ python3 tools/encoder/encode.py /tmp/fr out.dlx --profile sasi --preview p.png
|
||||
Two quality profiles ship from one codec and one decoder — `sasi` (110 KB/s) and
|
||||
`scsi` (280 KB/s) are two points on the same rate-distortion curve. Both are
|
||||
**ceilings**: lam is bisected per frame under a leaky bucket, so the profile's
|
||||
`lam` is a quality floor rather than a setting (`--fixed-lam` opts out). The codec is
|
||||
`lam` is a quality floor rather than a setting (`--fixed-lam` opts out).
|
||||
|
||||
There are **two** ceilings, on two different axes. The second is the 68000's
|
||||
decode budget: `mu` is bisected per frame against 833,333 cycles so the frame
|
||||
also *decodes* in time, which takes the worst sustained window from 37 frames
|
||||
over budget to 1 for 0.26 dB (FINDINGS 31). It is on by default; `--no-cpu-fit`
|
||||
restores session 7 behaviour. Unlike bytes, cycles have no bucket — there is no
|
||||
double buffer to decode ahead into, so it is a hard per-frame ceiling. The codec is
|
||||
a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four
|
||||
2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user