src/player/decode.s parses DLX1 and decodes straight into GVRAM. Verified pixel-exact over a 120-frame sequential run of the worst sustained window on the disc -- all four block modes, full temporal recursion, so the last frame is only right if all 120 were. In check.sh. It costs a mean of 81.7% of a 12fps frame budget, and 31% of frames exceed 100% (42% at scsi). CPU is now the binding constraint. FINDINGS 28. Three things that were believed and are not true: - The dual-display-path plan of FINDINGS 24.5/25.6 is incoherent. The compose path needs a RAM copy of the previous reconstruction; the direct path's selling point is that it keeps none. Mixing them shows stale pixels on 70 of 120 frames, worst frame 18.8% of the screen. Every coherent repair is dearer than not mixing, and 24.5's two figures were both copies with no decode in either, so there was never a crossover to find. One path ships, and the 96KB reference frame is gone. tools/analysis/10_pathmix_drift.py keeps the counterexample runnable; check.sh asserts it still reproduces. - The four block modes do not cost the same. V1 300, V4 448, RAW 400 cycles against the old model's flat 207.8. V4 is 25% of blocks and 50% of the cycles, and the mode decision charges it bytes it does not charge cycles for. tools/analysis/11_cpu_budget.py reproduces all four frames timed on the 68000 to within 1 point. Hand-derived timings agree to 0.5% on V1. - The container is big-endian but not aligned. Variable-length records laid end to end put frame 1's length field at an odd address, and move.l (a0)+ there is an address error: frame 0 decoded perfectly and then vectored into the IPL for 59 emulated seconds looking like a hang. Found by dumping PC, not by reading the source. Also: an all-V1 frame, the cheapest possible full redraw, is 110.5% of budget. No mode assignment fits a scene cut at 12fps. That one needs a decision, not a measurement. Next: charge cycles in the mode decision and bisect against 833,333 per frame, the way session 6 bisects lam against bytes -- but with no bucket, because a late frame cannot be banked. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
1374 lines
67 KiB
Markdown
1374 lines
67 KiB
Markdown
# Findings — session 1 (2026-08-23)
|
||
|
||
All numbers here are MEASURED unless marked ESTIMATE or FOLKLORE.
|
||
|
||
---
|
||
|
||
## 1. Source material
|
||
|
||
`DRAGONS_LAIR.iso` — 16 GB, UDF 2.x, **decrypted** (no AACS dir).
|
||
Loop-mounted read-only at `/media/reala-misaki/BDROM` via `udisksctl loop-setup -r -f`.
|
||
(7-Zip cannot read UDF 2.x; use the loop mount.)
|
||
|
||
- **224 `.m2ts` streams**, 1920x1080, **MPEG-2, progressive, 23.976 fps**
|
||
- Size histogram: 47 <5MB, 138 5-50MB, 22 50-150MB, 14 150-400MB, 3 >400MB
|
||
- The 185 sub-50MB streams are the **arcade branching scenes already split into
|
||
individual clips** — we get scene boundaries for free.
|
||
- Big streams are full-feature playthroughs: 00215 (1376s), 00216 (1151s), 00223 (566s)
|
||
- Typical scene clip ~60s (00203/00205/00199), some ~100s (00164/00212)
|
||
|
||
**Gotcha:** clip durations vary wildly. Always read `format=duration` and seek
|
||
relative to it. Seeking to a fixed offset silently yields 0 frames on short clips.
|
||
|
||
---
|
||
|
||
## 2. GVRAM layout [verified — see HARDWARE.md for source]
|
||
|
||
**One 16-bit word per pixel position in EVERY color mode.** Bit depth does not
|
||
change VRAM bandwidth; it only subdivides the word.
|
||
|
||
`addr = page_base + y*1024 + x*2` — adjacent pixels are 2 bytes apart in all modes.
|
||
|
||
Consequence: low bit depth buys **no speed**. 16-color mode is strictly worse than
|
||
256-color (same bus traffic, 1/16 the palette). Page-alias writes are hardware
|
||
auto-masked, so 16-color needs no software read-modify-write — but it's still
|
||
one word-access per pixel.
|
||
|
||
**Chosen: 256 colors, 256x192 active area.**
|
||
In 256-color mode P0=low byte, P1=high byte of each word. Sacrificing page 1 as a
|
||
double-buffer lets a `move.l` cover two pixel positions, enabling `movem.l` bursts
|
||
(12 regs = 48 bytes = 24 pixels). Identical blit cost to 65536-color mode but
|
||
**half the on-disk data**.
|
||
|
||
---
|
||
|
||
## 3. Content measurements (8 scenes sampled, 5s each at 40% into each clip)
|
||
|
||
| metric | mean | p90 |
|
||
|---|---|---|
|
||
| pixels changed / frame | 20.1% | 30.2% |
|
||
| **blit cost** | **~64k cycles** | **~97k cycles** |
|
||
| naive delta+RLE frame size | 15.5 KB | 19.6 KB |
|
||
|
||
Budget is **833,333 cycles/frame** @ 12fps on a 10MHz 68000.
|
||
|
||
### => THE CPU IS NOT THE BOTTLENECK. I/O IS.
|
||
Blit uses **under 8%** of budget. The naive row-span+RLE codec achieves only
|
||
**3.2:1**, giving **365 KB/s / 470 MB** at 24fps (~183 KB/s / 235 MB at 12fps).
|
||
|
||
Per-scene variance is extreme: static dialogue ~30 KB/s, action ~700 KB/s.
|
||
Any codec needs a hard bitrate ceiling, not just a good average.
|
||
|
||
### "Shot on twos" — ASSUMPTION FAILED
|
||
Dedupe found **zero** duplicate frames across all 8 scenes (`uniq=120/120`,
|
||
24.0 fps effective). This Blu-ray is a restoration where every frame is unique.
|
||
We do NOT get halved data for free. **Decimation to 12fps must be explicit.**
|
||
|
||
A weak alternation signature does exist (even-index pairs 40.7% vs odd 27.5%,
|
||
ratio 1.5x, with occasional true-duplicate pairs at 0.03-0.19%), but it is
|
||
irregular — Bluth mixed ones and twos; action is animated on ones.
|
||
|
||
---
|
||
|
||
## 4. MEASUREMENT TRAPS — read before trusting any pipeline number
|
||
|
||
Three separate false results were produced and caught this session. All three
|
||
looked plausible. Guard against them:
|
||
|
||
1. **Per-frame Floyd-Steinberg dithering destroys temporal coherence.**
|
||
Error diffusion is chaotic: a +/-1 input change cascades across the row and
|
||
produces a completely different index pattern. First run reported 31.5% pixels
|
||
changed with near-zero variance (median 31.6, p90 32.3, max 32.7) while source
|
||
mean-abs-diff was 0.09 — i.e. visually identical frames. That flat variance is
|
||
the tell: **real animation has scene-dependent variance; noise does not.**
|
||
Use no dithering (cel art is flat) or ordered/Bayer (spatially fixed, temporally stable).
|
||
|
||
2. **Temporal denoise smears motion.** `hqdn3d=4:3:6:4` — the `6:4` are temporal
|
||
params. It flattened real motion, which then measured as "no motion" and
|
||
produced an absurd 0.8 fps / 4 MB result. **Use spatial-only: `hqdn3d=4:3:0:0`.**
|
||
|
||
3. **Exact-match dedupe fails on a noisy source.** MPEG-2 grain means near-duplicate
|
||
frames differ by +/-1 and are never bit-exact. Use a threshold on
|
||
"% pixels differing by more than N levels", and pick the threshold from the
|
||
observed distribution, not a guess. A 2% threshold ate genuine animation when
|
||
mean consecutive change was only 0.9%.
|
||
|
||
**Sanity rule: if a result has suspiciously low variance, or is suspiciously
|
||
good, it is probably an artifact of the measurement, not a property of the content.**
|
||
|
||
Scripts kept in `tools/analysis/` — 01 and 02 are marked BROKEN deliberately as
|
||
regression references; 03 and 04 are the correct ones.
|
||
|
||
---
|
||
|
||
## 5. Storage interface — the SASI/SCSI split
|
||
|
||
> **SUPERSEDED IN PART.** The claim below that DMA means streaming "costs
|
||
> essentially no CPU" is **wrong** — see 19. The bandwidth figures here are
|
||
> folklore; the working figure is now **4 Mbps = 488 KB/s** (21).
|
||
|
||
[Yasuma, X68030 internal SCSI controller]
|
||
|
||
- Interface: **SCSI-1**, 50-pin, 5 MB/s bus spec
|
||
- Controller: **Fujitsu MB89352** SPC
|
||
- Transfer mode: **DMA** (via **HD63450** DMAC)
|
||
- Bus: X68000 original bus, **16-bit @ 10MHz**
|
||
|
||
**Even on the X68030, SCSI runs at 10MHz 16-bit DMA.** Storage bandwidth does
|
||
NOT scale with CPU — the controller sits on the original bus. HD63450's 12.5MHz
|
||
official ceiling is why the X68030 runs at 25MHz. An "HSCSI" TSR forces PIO/FIFO
|
||
transfer instead of DMA but was marginal even at 25MHz.
|
||
|
||
Because it's DMA, **streaming costs essentially no CPU** — this stacks with the
|
||
8% blit utilisation. The 68000 really is nearly idle.
|
||
|
||
### Model split — IMPORTANT
|
||
**The 10MHz models (original X68000, ACE, PRO, EXPERT) use SASI, not SCSI.**
|
||
Built-in SCSI starts at the X68000 **Super** (1990) and continues through XVI,
|
||
Compact, X68030. SCSI on earlier machines needs the **Sharp CZ-6BS1** board
|
||
in an I/O slot (MAME models this: `-exp1 cz6bs1`).
|
||
|
||
| target | bandwidth | naive codec (365 KB/s) | VQ codec (~30 KB/s) |
|
||
|---|---|---|---|
|
||
| SASI (stock ACE/EXPERT) | ~300-500 KB/s FOLKLORE | infeasible | comfortable |
|
||
| SCSI (Super+, or CZ-6BS1) | ~1 MB/s FOLKLORE | tight but viable | trivial |
|
||
|
||
Derived bounds (ESTIMATE): 16-bit @10MHz with 4-clock bus cycle = 5 MB/s absolute
|
||
ceiling; HD63450 single-address DMA ~8 clocks/word => ~2.5 MB/s practical ceiling,
|
||
before SCSI-1 async handshake and drive latency.
|
||
|
||
**No measured benchmark was obtained — see STATUS.md.** The ~300-500 KB/s and
|
||
~1 MB/s figures are folklore-grade; I could not find a primary measurement.
|
||
|
||
---
|
||
|
||
## 6. Codec decision: vector quantization (Cinepak-style)
|
||
|
||
Given ~8x CPU headroom and an I/O ceiling, spend CPU to buy bandwidth.
|
||
|
||
- Split frame into 4x4 blocks, encode each as a 1-byte index into a per-scene codebook
|
||
- Decode = 16-byte copy from a lookup table: nearly free
|
||
- A **full** frame = 256*192/16 = **3,072 bytes** — a hard 16:1 floor before delta
|
||
- Add block-level delta on top; action scenes ~2-3 KB/frame
|
||
- => roughly **30 KB/s, ~40 MB total**, with a *deterministic* bitrate ceiling
|
||
|
||
Divergence from the SNES project (below): use a **per-scene codebook with delta
|
||
updates**, not a per-frame rebuild. We trade adaptivity for bandwidth because we
|
||
have 2MB RAM to keep a codebook resident and CPU to spare.
|
||
|
||
**Risk not yet evaluated:** 4x4 VQ with a 256-entry codebook will visibly soften
|
||
detail. Bluth's fine ink linework is what suffers. Prototype and eyeball before committing.
|
||
|
||
---
|
||
|
||
## 7. Comparison: astrobleem/SNES-SuperDragonsLairArcade
|
||
|
||
Reached the **same core architecture independently** — "512 tiles per frame" is
|
||
vector quantization (8x8 codebook + tilemap). Good validation.
|
||
|
||
But: the SNES PPU has **no bitmap mode**, so tiles are forced on them by display
|
||
hardware. The X68000 has a real linear framebuffer, so VQ is a *compression
|
||
choice* we can tune or drop per-scene.
|
||
|
||
**MSU-1 is a bandwidth cheat we don't have.** It's a modern flash-cart coprocessor
|
||
giving memory-mapped streaming the real SNES never had. Their budget: 512 tiles x
|
||
32 bytes (4bpp 8x8) + tilemap ~= 18 KB/frame => **~430 KB/s** at 23.976fps.
|
||
That's *higher* than the 365 KB/s we'd reject on SASI. (ESTIMATE: my arithmetic on
|
||
their stated tile budget, not a measured figure.)
|
||
|
||
Where we're ahead: 256 simultaneous colors from a 65536 palette vs their 4bpp
|
||
sub-palettes needing a tile-aware palette optimizer plus a spatial smoothing pass
|
||
to hide 8x8 palette seams. That problem doesn't exist for us. Plus 68000@10MHz
|
||
vs 65816@3.58MHz, and 2MB vs 128KB.
|
||
|
||
**Most valuable thing in that repo is NOT the codec — it's `data/events/`:**
|
||
516 chapter definitions across 29 scenes as XML, plus
|
||
`data/chapter_event_inventory.md`. That's the arcade scene graph and input-timing
|
||
structure, entirely hardware-independent — the whole game-logic layer we'd
|
||
otherwise reverse-engineer from the arcade ROM.
|
||
|
||
**TODO: check their license before planning to reuse it.**
|
||
Their 516 chapters are finer-grained than our 224 Blu-ray streams, so mapping
|
||
their event table onto our footage means subdividing streams by timecode.
|
||
|
||
Caveat: all of the above is from README/repo-tree summaries, not their source.
|
||
|
||
---
|
||
---
|
||
|
||
# Findings — session 2 (2026-08-23)
|
||
|
||
## 8. CORRECTION to session 1: halving the framerate does NOT halve the bitrate
|
||
|
||
> **PARTLY SUPERSEDED.** The framerate correction stands. The
|
||
> "changed-spans + deflate = 247 KB/s" figure is a **compression upper bound,
|
||
> not a shippable design** — deflate decode does not fit the 68000's frame
|
||
> budget. See 17.2.
|
||
|
||
Session 1 measured 365 KB/s for naive delta+RLE at 24 fps and wrote
|
||
"(~183 KB/s at 12fps)". **That extrapolation is wrong.** Decimating to 12 fps
|
||
roughly doubles the per-frame delta, so the *rate* stays nearly flat.
|
||
|
||
Re-measured directly on 12 fps decimated frames (4 scenes, 66 frames):
|
||
|
||
| codec (all LOSSLESS w.r.t. the 256-colour frame) | B/frame | KB/s @12 | 22 min | ratio |
|
||
|---|---|---|---|---|
|
||
| raw 8bpp 256x192 | 49152 | 576 | 743 MB | 1.0:1 |
|
||
| session 1 row-span + RLE | 29055 | 340 | 439 MB | 1.7:1 |
|
||
| XOR vs prev + deflate | 30196 | 354 | 456 MB | 1.6:1 |
|
||
| **changed-spans + deflate** | **21110** | **247** | **319 MB** | **2.3:1** |
|
||
| changed-spans + LZMA | 18759 | 220 | 283 MB | 2.6:1 |
|
||
|
||
Session 1's own RLE re-measured at 12 fps gives **340 KB/s, not 183**.
|
||
Any plan that assumed 183 KB/s was based on a bad number.
|
||
|
||
Deflate-class entropy coding on top of the span payload is worth **1.4x** over
|
||
hand-rolled RLE, and LZ decode is cheap on a 68000 (byte copies), so the
|
||
lossless floor is ~247 KB/s / 319 MB. That is **infeasible on SASI** and
|
||
**tight but real on SCSI**.
|
||
|
||
## 9. Flat 4x4 VQ at k=256 is NOT acceptable — confirmed by eye
|
||
|
||
The risk flagged in 6 is real. At k=256, 4x4:
|
||
|
||
| scene | palette-only PSNR | after VQ | VQ loss |
|
||
|---|---|---|---|
|
||
| 00010 | 38.35 | 29.68 | 8.67 dB |
|
||
| 00020 | 39.90 | 32.67 | 7.22 dB |
|
||
| 00146 | 35.25 | 29.35 | 5.89 dB |
|
||
| 00181 | 41.92 | 32.87 | 9.05 dB |
|
||
|
||
Visually: Dirk's face disintegrates, teeth and eyes turn to mush, ink outlines
|
||
break into 4-pixel stair-steps, colour bleeds across block boundaries.
|
||
|
||

|
||
*Left: 1080p source. Middle: 256-colour palettised 256x192 — the quality ceiling,
|
||
and it is excellent. Right: flat 4x4 VQ at k=256. This is the result that killed
|
||
the flat-VQ architecture.*
|
||
|
||
**Crucially, the 256-colour palettised frame itself looks excellent.** Flat cel
|
||
art with a per-scene median-cut palette and no dithering is near-transparent
|
||
(35-42 dB). So the palette is not the problem and 256 colours is not the
|
||
problem — **block VQ is**. The quality ceiling we should hold ourselves to is
|
||
the palettised frame, not the 1080p source.
|
||
|
||
## 10. Hybrid VQ (Cinepak V1/V4 + SKIP) — this is the codec
|
||
|
||
Per 4x4 block, choose by rate-distortion: SKIP (reuse previous frame),
|
||
V1 (one 4x4 codeword, 1 byte), or V4 (four 2x2 codewords, 4 bytes),
|
||
with a 2-bit-per-block mode header. `lam` is the lagrangian rate knob.
|
||
|
||
Measured, k1=k4=256, 4 scenes (mean of the per-scene table in the session log):
|
||
|
||
| lam | PSNR | loss vs palette | SKIP% | V1% | V4% | B/frame | KB/s @12 |
|
||
|---|---|---|---|---|---|---|---|
|
||
| 0 (max quality) | 33.9 | 4.9 | 30.8 | 18.5 | 50.8 | 7574 | 88.8 |
|
||
| 200 | 31.9 | 5.9 | 44.0 | 37.6 | 18.4 | 4183 | 49.0 |
|
||
| 1000 | 31.6 | 7.3 | 47.4 | 47.7 | 4.9 | 2841 | 33.3 |
|
||
| 5000 | 25.5 | 13.3 | 55.6 | 44.4 | 0.0 | 2134 | 25.0 |
|
||
|
||
At a **matched ~30 KB/s** the hybrid beats flat 4x4 VQ by ~1 dB, and unlike flat
|
||
VQ it keeps scaling: at 89 KB/s it reaches within **4.9 dB of the palette
|
||
ceiling**, which flat VQ cannot reach at any bitrate.
|
||
|
||
Note V4% collapses to 0 at lam=5000 — that is the knob doing exactly what it
|
||
should: under a hard ceiling, detail blocks are the first thing sacrificed.
|
||
|
||
## 11. Codebook size sweep (flat 4x4, for reference)
|
||
|
||
> **SUPERSEDED.** The k=1024 result below is an artifact of a rate model that
|
||
> charged 1 byte for a 10-bit index. k=256 ships. See 14.
|
||
|
||
| block | k | PSNR | loss | key B | changed% | KB/s @12 | codebook RAM |
|
||
|---|---|---|---|---|---|---|---|
|
||
| 4x4 | 256 | 30.46 | 8.39 | 3072 | 52.7 | 28.5 | 8K |
|
||
| 4x4 | 1024 | 32.89 | 5.96 | 3840 | 56.6 | 35.6 | 32K |
|
||
|
||
+2.4 dB for 24K more RAM and 7 KB/s. With 2 MB of RAM, a 1024-entry codebook is
|
||
cheap and clearly worth it. (RAM figure is the word-expanded form the blitter
|
||
wants: k * 16 px * 2 bytes.)
|
||
|
||
## 12. Source framing — OPEN
|
||
|
||
The Blu-ray is **full-frame 1920x1080 16:9 with no pillarboxing**. The arcade
|
||
original is 4:3. The extractor currently centre-crops 1440x1080, which is the
|
||
arcade-faithful choice but discards image the 2006 remaster added. Options are
|
||
`crop` (default), `squash`, `wide` in `tools/encoder/extract.py`.
|
||
**Not yet decided; needs an eyeball comparison against arcade reference.**
|
||
|
||
## 13. Stream inventory correction
|
||
|
||
Session 1 said "typical scene clip ~60s". Sampled directly: the ~3-5 MB streams
|
||
are **1.2-1.7 s** clips — these are the individual arcade death/action moments,
|
||
which is exactly the granularity the game logic needs. Some 60 s streams
|
||
(e.g. 00203) are **menu screens, not content**. Any survey must classify
|
||
menu vs content before averaging, or the bitrate numbers are diluted by static
|
||
menus.
|
||
|
||
## 14. A FOURTH false-good result — and the correction
|
||
|
||
Add this to the 4 list. The mechanism was new but the shape was identical.
|
||
|
||
**The false result:** flat and hybrid VQ both showed **+2.4 dB for k=1024 over
|
||
k=256** at an apparently similar bitrate, which made a 1024-entry codebook look
|
||
like an obvious win. The k=1024 quality ladder rendered from that run looked
|
||
great at "45 KB/s".
|
||
|
||
**The bug:** the rate-distortion model in `vq_hybrid.encode()` charged **1 byte**
|
||
per codebook index unconditionally. A 1024-entry codebook needs a **10-bit index,
|
||
stored as 2 bytes**. So every k=1024 measurement understated the V1 and V4
|
||
payload by exactly 2x, *and* the lagrangian mode decision was choosing V4 on the
|
||
belief that four codewords cost 4 bytes when they cost 8.
|
||
|
||
**After charging the true index cost** (`idx_bytes` is now explicit and defaults
|
||
from the codebook size), matched-bitrate comparison on scene 00020:
|
||
|
||
| KB/s | k=256 (1-byte idx) | k=1024 (2-byte idx) |
|
||
|---|---|---|
|
||
| ~32-42 | **33.87 dB** @ 32.5 | 28.91 dB @ 42.3 |
|
||
| ~44-52 | **34.80 dB** @ 44.1 | 35.13 dB @ 52.5 |
|
||
| ~72-86 | **35.87 dB** @ 72.2 | 36.51 dB @ 86.0 |
|
||
|
||
k=1024 buys +0.3 to +0.6 dB for +19% bitrate — a wash at best — and at the low
|
||
end where the SASI profile lives it is **5 dB worse**, because the 2-byte index
|
||
floor dominates once V4 is priced out.
|
||
|
||
**k=256 with 1-byte indices is the shipping choice.** It is also the better
|
||
decoder: a plain `move.b` index with no alignment case, and an 8 KB codebook
|
||
instead of 32 KB.
|
||
|
||
**The general lesson, again:** the comparison was not wrong about VQ, it was
|
||
wrong about *cost*. When a knob looks like a free win, check that the rate model
|
||
is charging for it. Same failure family as 4.1-4.3: a plausible number produced
|
||
by a pipeline that was not measuring what it claimed to measure.
|
||
|
||
## 15. Rate-distortion curve of the shipping codec (k=256, corrected)
|
||
|
||
Scene 00020 (Dirk screaming, close-up face — the hardest case for linework),
|
||
and 00146. Includes the 2-bit-per-block mode header. No entropy coding yet.
|
||
|
||
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | SKIP | V1 | V4 | RAW |
|
||
|---|---|---|---|---|---|---|---|---|
|
||
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~37% | ~24% | ~13% | ~26% |
|
||
| 100 | 35.87 | 72.2 | 29.04 | 72.5 | ~41% | ~34% | ~21% | ~4% |
|
||
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | ~44% | ~42% | ~14% | 0% |
|
||
| 800 | 33.87 | 32.5 | 27.77 | 36.1 | ~46% | ~48% | ~5% | 0% |
|
||
| 2000 | 27.57 | 25.5 | 24.88 | 30.2 | ~50% | ~49% | ~1% | 0% |
|
||
|
||
Palette ceilings: 00020 = 39.90 dB, 00146 = 35.25 dB.
|
||
|
||

|
||
*The shipping codec across the rate knob. Top: source, palette ceiling, lam=25.
|
||
Bottom: lam=100 (`scsi` profile), lam=300 (`sasi` profile), lam=800.
|
||
Both shipping profiles hold Bluth's linework; the failure only starts past lam=800.*
|
||
|
||
Two things to read off this table:
|
||
- **The cliff is between lam=800 and lam=2000.** That is where V4 is priced out
|
||
entirely and detail blocks have nowhere to go. Do not ship past lam~800.
|
||
- **RAW is doing real work at high bitrate** (26% of blocks at lam=25) and
|
||
vanishes by lam=300. It is what makes the top of the curve reach the palette
|
||
ceiling, and it costs the decoder nothing — RAW is the cheapest mode to blit.
|
||
|
||
## 16. Licences cleared for the game-logic layer
|
||
|
||
Both checked this session:
|
||
|
||
- **astrobleem/SNES-SuperDragonsLairArcade — MIT**, "Copyright (c) 2026 Chad
|
||
Doebelin". `data/events/` holds 516 XML chapter definitions with timing and
|
||
event data. Reusable with attribution.
|
||
- **icculus/DirkSimple — zlib.** Independent from-scratch reimplementation of
|
||
the game logic in Lua, scene/timing tables in `game.lua`. Also permissive.
|
||
|
||
Having **two independent permissively-licensed transcriptions** of the arcade
|
||
scene graph is better than one: they can be diffed against each other to catch
|
||
transcription errors before any of it is committed to 68000 tables.
|
||
|
||
## 17. The profiles were set far too low — and entropy coding is a CPU trap
|
||
|
||
> **PARTLY SUPERSEDED.** 17.1's diagnosis (the profiles were not derived from
|
||
> hardware) and 17.2's CPU analysis both stand. But 17 reasoned against a
|
||
> misread bandwidth of 4 MB/s; the correct figure is **4 Mbps = 488 KB/s**, so
|
||
> the "ship pixel-exact if SCSI sustains >=800 KB/s" conclusion in 17.5 is
|
||
> **not available**. See 18 and 21.
|
||
|
||
Prompted by the user asking why the SCSI profile was only 75 KB/s. It should not
|
||
have been. Two separate errors, one of them serious.
|
||
|
||
### 17.1 The profile bitrates were not derived from the hardware at all
|
||
|
||
They were read off the knee of the rate-distortion curve and then presented as
|
||
though bandwidth-driven. Against the (folklore) bus figures from 5:
|
||
|
||
| profile | was | bus figure | utilisation |
|
||
|---|---|---|---|
|
||
| `sasi` | 45 KB/s | ~300-500 KB/s | **12%** |
|
||
| `scsi` | 75 KB/s | ~1 MB/s | **7%** |
|
||
|
||
Nothing justified leaving 90% of the pipe unused. Raised to `sasi` 110 KB/s
|
||
(lam=60) and `scsi` 280 KB/s (lam=10), which is 35% and 28% utilisation —
|
||
still conservative, because the bus figures are folklore.
|
||
|
||
### 17.2 CPU is NOT the reason to stay low — but entropy coding would be
|
||
|
||
Budget is 833,333 cycles/frame at 12 fps. At session 1's measured ~6.5 cycles
|
||
per GVRAM pixel write:
|
||
|
||
| work | cycles | % of budget |
|
||
|---|---|---|
|
||
| blit 20.1% of pixels (session 1's 24fps figure) | 64k | 7.7% |
|
||
| blit 40% of pixels (the same content at 12fps) | 128k | 15.3% |
|
||
| **blit the FULL frame, every frame** | **319k** | **38.3%** |
|
||
| deflate decode, ~30 KB/frame output | 1,800k | **216%** |
|
||
| LZ4/LZSS decode, ~30 KB/frame output | 450k | **54%** |
|
||
|
||
Two conclusions, and the second one corrects 8:
|
||
|
||
- **Raising the VQ bitrate is nearly free on CPU.** Even a full-frame pixel-exact
|
||
blit fits in 38% of budget, and VQ decode is table copies — RAW, the mode that
|
||
dominates at high bitrate, is the *cheapest* mode to blit, not the dearest.
|
||
- **The 247 KB/s "lossless changed-spans + deflate" figure in 8 is a compression
|
||
upper bound, NOT a shippable design.** Deflate's Huffman decode is bitwise and
|
||
costs about 2.2x the entire frame budget on a 68000. Even byte-oriented LZ4 at
|
||
54% leaves too little beside a 38% blit. **Do not plan on entropy coding.**
|
||
All profile bitrates are raw payload.
|
||
|
||
This inverts session 1's "the CPU is idle, I/O is the ceiling" for the *decode*
|
||
path specifically: the blit is cheap, but any bit-oriented decompressor is not.
|
||
VQ is the right architecture precisely because its decode is a table copy.
|
||
|
||
### 17.3 The hybrid at lam=0 IS the lossless codec
|
||
|
||
Measured, un-entropy-coded raw payload, and deflated for reference only:
|
||
|
||
| scene | lam=0 raw | lam=0 deflated | lossless changed-spans+deflate | PSNR |
|
||
|---|---|---|---|---|
|
||
| 00020 | 442.1 KB/s | 274.5 KB/s | 267.3 KB/s | 39.90 = ceiling |
|
||
| 00146 | 467.6 KB/s | 223.2 KB/s | 219.1 KB/s | 35.25 = ceiling |
|
||
|
||
The hybrid at `lam=0` converges to within 3% of the purpose-built lossless coder.
|
||
That confirms the architecture unifies: there is no separate lossless path to
|
||
maintain, just the same bitstream with the knob open.
|
||
|
||
### 17.4 Full curve in raw (shippable) bytes
|
||
|
||
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | RAW% |
|
||
|---|---|---|---|---|---|
|
||
| 0 | **39.90** (exact) | 442.1 | **35.25** (exact) | 467.6 | ~76% |
|
||
| 10 | 39.38 | 248.1 | 32.27 | 305.2 | ~44% |
|
||
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~26% |
|
||
| 60 | 36.94 | 108.0 | 29.61 | 103.1 | ~10% |
|
||
| 150 | 35.31 | 55.6 | 28.63 | 56.1 | ~1% |
|
||
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | 0% |
|
||
|
||
### 17.5 This makes the blocked disk benchmark critical-path
|
||
|
||
Session 1 judged it "NOT on the critical path" because VQ at 30 KB/s was correct
|
||
whether SASI did 300 or 600 KB/s. That reasoning no longer holds. The profiles
|
||
now sit at 110 and 280 KB/s, close enough to the folklore ceilings that the
|
||
error bars matter, and **if SCSI sustains >=800 KB/s the correct `scsi` profile
|
||
is lam=0 — pixel-exact video.** Whether this port ships transparent or lossy on
|
||
SCSI is now waiting on one measurement.
|
||
|
||
## 18. Peak-to-mean burstiness — the mean was hiding the problem
|
||
|
||
> **SUPERSEDED — DO NOT ACT ON THIS SECTION.** The peak-vs-sustained comparison
|
||
> below is the **wrong test**. With a ring buffer the correct test is cumulative
|
||
> demand vs cumulative supply, and both profiles pass it with **zero required
|
||
> prefill**. `scsi` at lam=10 ships. See 21. The per-frame peak numbers
|
||
> themselves are still valid data; only the conclusion drawn from them is wrong.
|
||
|
||
Prompted by the user clarifying that the bandwidth figure is **4 Mbps = 488 KB/s**,
|
||
not 4 MB/s. That is ~8x tighter than what 17 was reasoning against, and it
|
||
changes the answer.
|
||
|
||
Per-frame instantaneous rate (video + 7.8 KB/s audio), 12 fps:
|
||
|
||
| scene | lam | mean | p90 | **max** | peak/mean | max as % of 488 KB/s |
|
||
|---|---|---|---|---|---|---|
|
||
| 00010 | 60 | 95.0 | 127.3 | 138.8 | 1.46 | 28.4% |
|
||
| 00010 | 10 | 198.9 | 266.1 | 284.0 | 1.43 | 58.2% |
|
||
| 00020 | 60 | 115.8 | 155.4 | 222.3 | 1.92 | 45.5% |
|
||
| 00020 | 10 | 255.9 | 391.2 | **470.8** | 1.84 | **96.4%** |
|
||
|
||
**The `scsi` profile as committed in f0f2f80 does not fit 4 Mbps.** Its mean is a
|
||
comfortable 52% of the pipe, but it peaks at 96.4% — and a frame that arrives
|
||
late is a *dropped frame*, not a slow one. Sizing a real-time stream on the mean
|
||
is the mistake; peak/mean is 1.4-1.9x on 1.2-1.7s clips and will be worse across
|
||
a full scene.
|
||
|
||
Two ways out, and only one is good:
|
||
- Size for the peak: `lam=25`, mean 194 KB/s. Costs a full step of quality.
|
||
- **Rate-control to the mean and carry a leaky bucket:** `lam=10` fits, and buys
|
||
back +0.7 dB (00020) / +1.2 dB (00146).
|
||
|
||
`ratectl.py` was written in session 2 but **never wired into `encode.py`**. This
|
||
demotes that from a loose end to the highest-value unfinished work in the repo.
|
||
|
||
## 19. Cycle-stealing DMA is not free DMA — 5 was wrong
|
||
|
||
FINDINGS 5 concluded "because it's DMA, streaming costs essentially no CPU —
|
||
this stacks with the 8% blit utilisation. The 68000 really is nearly idle."
|
||
|
||
The HD63450 steals bus cycles from the 68000 at roughly 8 clocks per 16-bit word:
|
||
|
||
| stream | words/s | clocks/s | CPU stolen | + full-frame blit |
|
||
|---|---|---|---|---|
|
||
| 110 KB/s | 56,320 | 450,560 | 4.5% | 42.8% |
|
||
| 250 KB/s | 128,000 | 1,024,000 | 10.2% | 48.5% |
|
||
| 450 KB/s | 230,400 | 1,843,200 | 18.4% | 56.7% |
|
||
| 488 KB/s | 249,856 | 1,998,848 | 20.0% | 58.3% |
|
||
|
||
At the rates the profiles now use, streaming costs **10-20% of the machine**.
|
||
Still affordable — nothing here breaks — but **bandwidth and CPU are one budget,
|
||
not two**, and any future headroom argument has to spend from both. The
|
||
"nearly idle" framing should not be reused.
|
||
|
||
(The 8 clocks/word figure is session 1's ESTIMATE from HD63450 timing, not a
|
||
measurement. It is the weakest link in this table.)
|
||
|
||
## 20. Where the profiles should come from
|
||
|
||
`tools/encoder/profile_gen.py` now derives lam from a bandwidth figure rather
|
||
than from the shape of the RD curve, accounting for audio, peak/mean, and
|
||
reporting DMA steal. Full benchmarking methodology — and why MAME cannot answer
|
||
the bandwidth question — is in `docs/BENCHMARK.md`.
|
||
|
||
The 4 Mbps figure itself is **user-supplied and its provenance is not recorded**.
|
||
Every profile now hangs off it, so it is worth pinning down.
|
||
|
||
## 21. Correction to 18 — the peak test was the wrong test
|
||
|
||
18 flagged that `scsi` "does not fit 4 Mbps" because a frame peaked at 96.4% of
|
||
the sustained rate. **That was the wrong comparison**, and the user was right to
|
||
push back. It measured instantaneous frame demand against a sustained rate as if
|
||
they had to match frame-by-frame. They do not: the disk keeps filling *during*
|
||
the frame, and any shortfall is absorbed by a ring buffer.
|
||
|
||
The correct test is whether **cumulative** demand ever outruns cumulative supply.
|
||
Simulated at a constant 488 KB/s fill, 12 fps, using the real per-frame sizes:
|
||
|
||
| scene | lam | mean KB/s | worst frame | **required prefill** | stall tolerance @256KB |
|
||
|---|---|---|---|---|---|
|
||
| 00010 | 10 | 198.9 | 23.67 KB | **0.0 KB** | 15.4 frames |
|
||
| 00020 | 10 | 255.9 | 39.23 KB | **0.0 KB** | 12.0 frames |
|
||
| 00146 | 10 | **313.0** | 42.10 KB | **0.0 KB** | 9.8 frames |
|
||
| 00181 | 10 | 211.1 | 25.25 KB | **0.0 KB** | 14.6 frames |
|
||
| (all) | 60 | 95-116 | 11-19 KB | **0.0 KB** | 26-32 frames |
|
||
|
||
Fill delivers **40.69 KB per frame time**. Only one measured frame exceeds that
|
||
(00146, 42.10 KB) and it is recovered by the following frame. **No scene needs
|
||
any prefill at all**, and a 256 KB buffer — 12.5% of RAM — carries ~1 second of
|
||
stall tolerance at `lam=10`, which is orders of magnitude more than an SD-backed
|
||
seek requires.
|
||
|
||
`scsi` at `lam=10` stands. The hardest scene sampled (00146) runs 313 KB/s mean,
|
||
64% of the pipe, with zero underrun risk.
|
||
|
||
### Why SD-backed changes the sizing rule
|
||
The deployment target is BlueSCSI / SCSI2SD, not a period spinning drive. That
|
||
was noted as a caveat in 5 but its consequence was not carried through:
|
||
|
||
- The sustained rate is a **bus-limited constant**, not an average over variable
|
||
seek latency. There is no long tail to leave margin for.
|
||
- Seek is ~microseconds, so branch-point stalls are a non-issue against a buffer
|
||
measured in whole seconds.
|
||
- Therefore we can size much closer to the ceiling than spinning-disk practice
|
||
would suggest. Conservative margins here are cargo-culted from a constraint
|
||
this deployment does not have.
|
||
|
||
**The SASI/SCSI split is about BUS PROTOCOL, not media.** SD emulation removes
|
||
seek latency from both, but a SASI bus is still slower than a SCSI one. Two
|
||
profiles remain the right design; both are now predictable constants rather than
|
||
distributions.
|
||
|
||
### What rate control is actually for now
|
||
Its value drops from load-bearing to **insurance**. Intra-scene peaks are a
|
||
non-problem. But we have measured **4 clips of 1.2-1.7s** out of 224 streams, and
|
||
00146 already runs 23% hotter than 00020. A sustained action sequence could
|
||
plausibly exceed the pipe where a 1.7s clip does not. Rate control gives a
|
||
*deterministic ceiling* across content we have not measured yet — which was the
|
||
original reason for choosing VQ over a lossless delta in the first place.
|
||
|
||
Still worth wiring in. No longer a blocker for shipping `scsi` at `lam=10`.
|
||
|
||
## 22. The display path, measured — first real frame on the X68000
|
||
|
||
Everything before this section was Python-side or a headless `-video none` run.
|
||
This is the first time pixels reached an emulated X68000 screen, and it produced
|
||
four hardware facts and one blocker that no amount of reasoning would have found.
|
||
|
||
**Scope — read this before quoting the result.** The X68000's *video* hardware
|
||
did the rendering: CRTC, GVRAM page decoding and the 15-bit+I palette lookup are
|
||
all genuinely emulated, which is why the output is bit-exact against the
|
||
hardware's colour math. But the pixels were written into GVRAM by a MAME Lua
|
||
script calling `SP:write_u16()` — the host poking emulated memory. **No 68000
|
||
instruction was executed to draw this frame.**
|
||
|
||
The equivalent is proving a framebuffer works by writing to it from a debugger.
|
||
It says the display path is correct; it says *nothing* about whether the 68000
|
||
can fill that framebuffer in time. Lua writes cost zero 68000 cycles, so the 38%
|
||
full-frame blit estimate that the entire CPU budget rests on remains completely
|
||
unvalidated. That is next step (2), the decoder skeleton, and it is untouched.
|
||
|
||
Reproduce:
|
||
```
|
||
python3 tools/bench/prep_frame.py <framedir> tmp/frame.bin 0
|
||
cd tmp && SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
|
||
-sound none -nothrottle -plugins -autoboot_script ../tools/bench/show_frame.lua \
|
||
-snapshot_directory ./snap -snapview native -seconds_to_run 6
|
||
```
|
||
|
||
### 22.1 The blocker: CRTC R20 bit 11 hides the graphics layer
|
||
The IPL leaves **CRTC R20 (`$E80028`) = `0x0B16`**. Bit 11 is *"G-VRAM set to
|
||
buffer"*, and MAME's `x68k_v.cpp` bails out of `draw_gfx()` on it outright:
|
||
|
||
```c
|
||
if (m_crtc->gfx_layer_buffer()) // if graphic layers are set to buffer, they aren't visible
|
||
return false;
|
||
// x68k_crtc.h: bool gfx_layer_buffer() const { return BIT(m_reg[20], 11); }
|
||
```
|
||
|
||
While that bit is set, GVRAM writes still land and read back correctly — which
|
||
is exactly what makes it so misleading. Six separate attempts at the video
|
||
controller (`$E82400/$E82500/$E82600`) rendered black with every register
|
||
reading back the intended value. **The video controller was never the problem.**
|
||
|
||
`R20` bits 9-8 select the colour setup, and this determines how `$C00000` is
|
||
decoded: `0x0300` = 65536c (16 bits/word), `0x0100` = 256c (low byte),
|
||
`0x0000` = 16c (4 bits). Set `R20 = 0x0116` for our mode.
|
||
|
||
### 22.2 Monitor contrast: the IPL leaves it at 14, not 15
|
||
`$E8E001` bits 3-0 are monitor contrast; MAME does
|
||
`m_screen->set_brightness(contrast * 0x11)`. The IPL leaves it at **14**, which
|
||
scales all output to 14/15 = 93.3%. Every rendered colour came out ~7% dark
|
||
until this was set to 15. **The player must write `$E8E001 = 15` at startup.**
|
||
|
||
Contrast `0` blanks the screen entirely (`x68k_v.cpp:661`) — that is the cheap
|
||
fade-to-black for scene transitions, no palette animation required.
|
||
|
||
### 22.3 Palette format CONFIRMED (was previously an assumption)
|
||
`PALETTE(config, m_gfxpalette).set_format(2, &x68k_state::GGGGGRRRRRBBBBBI, 256)`
|
||
|
||
```
|
||
bit 15..11 10..6 5..1 0
|
||
GGGGG RRRRR BBBBB I <- I is a shared LSB for all three channels
|
||
```
|
||
Expansion is `pal6bit((field << 1) | I)`, i.e. `(v << 2) | (v >> 4)`.
|
||
With contrast at 15, **all 256 entries render exactly as this predicts** — the
|
||
frame is pixel-identical, not merely close. GVRAM line stride is confirmed as
|
||
512 words = 1024 bytes, matching `HARDWARE.md`.
|
||
|
||
### 22.4 A new quality ceiling: the 15-bit palette costs 38.88 dB
|
||
> **Superseded by 23.3.** The 38.88 dB figure assumed the shared LSB `I` is
|
||
> always 1. Choosing `I` per palette entry by minimum error lifts the ceiling to
|
||
> **40.81 dB** on the same frame. The conclusion below ("`scsi` is close to
|
||
> display-transparent") is therefore weaker than stated — there is ~2 dB more
|
||
> headroom than this section claims.
|
||
Section 3 called the 256-colour palettised frame "the real quality ceiling".
|
||
That was measured in 24-bit RGB. The hardware palette only stores 5 bits per
|
||
channel plus a shared LSB, so there is a **second** quantisation below it:
|
||
|
||
| stage | PSNR |
|
||
|---|---|
|
||
| 24-bit palettised source -> X68000 15-bit+I display | **38.88 dB** |
|
||
| `scsi` profile codec error (00020, FINDINGS 15) | 39.4 dB |
|
||
|
||
The codec's error at `scsi` is **the same order as the display's own error**.
|
||
On real hardware `scsi` is therefore close to display-transparent, and pushing
|
||
`lam` below 10 buys quality the monitor cannot show. This bounds how much the
|
||
`scsi` profile is worth raising — it does not change the profiles themselves.
|
||
|
||
Caveat: measured on one frame (00020 f0001). It is a property of the palette,
|
||
not the content, so it should generalise, but it has not been checked across
|
||
scenes.
|
||
|
||
### 22.5 Why the first frame appears twice
|
||
GVRAM is a 512-pixel-wide page while the IPL's CRTC is still in its 768-wide
|
||
text timing, so the layer repeats at exactly x=512. This is correct hardware
|
||
behaviour, not a bug. The player sets its own CRTC mode and the wrap disappears.
|
||
No CRTC timing table has been written yet — the harness deliberately keeps the
|
||
IPL's timing so that no invented CRTC values are in play.
|
||
|
||
|
||
## 23. A real CRTC mode: 256x192 inside 256x256 (session 4)
|
||
|
||
Session 3's harness borrowed the IPL's 768x512 text timing and invented no CRTC
|
||
values, which is why the frame repeated at x=512 (22.5). This session derived a
|
||
real 256x256 mode table from MAME 0.277 source and verified it by snapshot.
|
||
Table: `tools/bench/crtc_mode.lua`. Regression test: `tools/bench/verify_frame256.py`.
|
||
|
||

|
||
|
||
*Left: palettised source. Right: the emulated X68000's native 256x512 raster —
|
||
256 dots wide, 512 scanlines carrying 256 double-scanned graphics rows, with the
|
||
192-row picture letterboxed in true black.*
|
||
|
||
### 23.1 The table, and why it needed no guessing
|
||
`refresh_mode()` in `x68k_crtc.cpp` selects the dot clock as
|
||
`(reg20 bit4 ? 69.55199MHz : 38.86363MHz) / div`, with `div` from a ladder keyed
|
||
on `reg20 & 0x1f`. Three entries matter:
|
||
|
||
| `reg20 & 0x1f` | div | dot clock | mode |
|
||
|---|---|---|---|
|
||
| `0x16` | 2 | 34.776 MHz | IPL's 768 wide, 31.5kHz |
|
||
| `0x11` | 3 | 23.184 MHz | 512 wide, 31.5kHz |
|
||
| `0x10` | 6 | 11.592 MHz | **256 wide, 31.5kHz, graphics double-scanned** |
|
||
|
||
The IPL's `R00 = 137` gives `m_htotal = (137+1)*8 = 1104` dots, and
|
||
`34.776e6 / 1104 = 31500.0 Hz` **exactly**. Holding the same line rate at div 6
|
||
needs `11.592e6 / 31500 = 368` dots `= 46` chars, so `R00 = 45`.
|
||
|
||
`368 = 1104/3` exactly, so every horizontal register is the 768-mode value
|
||
divided by three, and the active window divides without remainder:
|
||
`(124-28)/3 = 32` chars `= 256` dots. **No horizontal value was recalled or
|
||
estimated.** Only the blanking split rounds: the 768 mode is sync/back/front =
|
||
14/14/14 chars, `/3 = 4.67` each, and the closest integer triple summing to
|
||
`46-32 = 14` is 5/5/4.
|
||
|
||
| reg | value | meaning |
|
||
|---|---|---|
|
||
| R00 | 45 | H total, 46 chars = 368 dots -> 31500.0 Hz |
|
||
| R01 | 5 | H sync end (3.45 us) |
|
||
| R02 | 10 | H display begin -> `hbegin = 81` |
|
||
| R03 | 42 | H display end -> `hend = 336`; inclusive width `336-81+1 = 256` |
|
||
| R04 | 567 | V total, 568 scanlines -> 55.46 Hz |
|
||
| R05 | 5 | V sync end |
|
||
| R06 | 40 | V display begin -> `vbegin = 41` |
|
||
| R07 | 552 | V display end -> 512 scanlines = 256 double-scanned rows |
|
||
| R08 | 27 | H sync adjust (MAME stores it and never reads it) |
|
||
| R20 | `0x0110` | display (not buffer), 256-colour, 31.5kHz, 256x256 |
|
||
|
||
**The vertical registers are NOT halved**, which is the one thing that looks
|
||
wrong and is not. The CRTC still generates a 568-line raster; "256 lines" is a
|
||
graphics-layer double-scan applied in `draw_gfx()` (`x68k_v.cpp:401`), not a
|
||
change to the raster. Halving R04 would ask the monitor for 110 Hz. MAME emits
|
||
a `visarea larger then reg[20]` logerror for this; it is cosmetic.
|
||
|
||
Total blanking time is identical to the 768 mode (112 dots at 11.592 MHz =
|
||
336 dots at 34.776 MHz = 9.66 us), which is the property a real monitor cares
|
||
about — so this table should be safe on hardware, though that is untested.
|
||
|
||
### 23.2 MAME's double-scan is phase-shifted by one raster line
|
||
`get_gfx_pixel()` indexes `m_gfxbitmap.pix(scanline / divisor, pixel)` using the
|
||
**absolute** scanline, and `vbegin = 41` is odd. So in the native 256x512
|
||
snapshot the identical row pairs are `(1,2), (3,4), ...` and row 0 is a lone
|
||
half-line. Even rows are graphics rows 0..255. This cost a false failure before
|
||
it was understood; the regression test now asserts the shifted pairing
|
||
explicitly so a change in MAME's behaviour is visible rather than confusing.
|
||
|
||
### 23.3 The shared LSB `I` must be chosen per palette entry — worth 1.96 dB
|
||
Session 3's `pack()` hardcoded `I = 1`. That is not free: `I` is shared by all
|
||
three channels and each renders as `pal6bit((field << 1) | I)`, so with `I = 1`
|
||
the darkest reachable value is `pal6bit(1) = 4`, and **true black does not
|
||
exist**. Choosing `I` per entry to minimise summed squared error over R,G,B:
|
||
|
||
| rule | ceiling vs 24-bit palettised (00020 f0001) | entries with I=0 |
|
||
|---|---|---|
|
||
| `I = 1` fixed (session 3) | 38.85 dB | 0 |
|
||
| `I` per entry, min squared error | **40.81 dB** | 102 / 256 |
|
||
|
||
Nearly **2 dB for free**, and 102 of 256 entries want `I = 0` — this is not a
|
||
corner case. It supersedes the ceiling in 22.4 and means `scsi` has about 2 dB
|
||
more headroom before it hits the display than that section claimed.
|
||
|
||
The encoder does not yet do this. `tools/encoder/` still emits 24-bit palettes
|
||
and the packing happens Lua-side; whatever eventually writes X68000 palette
|
||
words must use the per-entry rule.
|
||
|
||
### 23.4 Letterboxing requires a reserved black palette entry
|
||
GVRAM cleared to zero displays **palette entry 0**, and a free mediancut palette
|
||
puts a real image colour there — on 00020 f0001 it was `(206,192,176)`, used by
|
||
210 image pixels, so it cannot simply be repurposed. A 256x192 picture in a
|
||
256x256 mode has 64 blank rows, so the palette must be built with **255 colours
|
||
plus a reserved black at index 0** (`prep_frame.py --reserve-black`). Combined
|
||
with 23.3, entry 0 also needs `I = 0` or the bars sit at RGB (4,4,4).
|
||
|
||
Cost: one of 256 entries. Measured quality effect: none visible — the ceiling
|
||
figure in 23.3 is already measured on the 255-colour palette.
|
||
|
||
### 23.5 What is still not proven
|
||
GVRAM was again filled from Lua. **No 68000 instruction has drawn a pixel yet**,
|
||
and the 38% full-frame blit estimate underpinning the CPU budget remains
|
||
unvalidated. What this section adds is that the *target mode* is now real, so
|
||
68000 code has a defined geometry to write into: 256 words per visible row, a
|
||
1024-byte line stride, and rows 32..223 of a 256-row page.
|
||
|
||
---
|
||
|
||
## 24. The blit, measured on the 68000 — the 38% estimate was wrong (session 5)
|
||
|
||
**The first 68000 instructions in this project to draw a pixel.** Everything in
|
||
22 and 23 was GVRAM filled from Lua, which costs zero 68000 cycles. This section
|
||
replaces the estimate that the whole CPU budget rested on with a measurement.
|
||
|
||
Harness: `tools/bench/blit.s` + `tools/bench/blit.lua`. Four variants of a
|
||
full-frame 256x192 paint, each looped to run ~4 emulated seconds, timed from
|
||
`machine.time` between two flag writes by the 68000 itself.
|
||
|
||
| variant | what it does | cycles/frame | % of a 12fps frame |
|
||
|---|---|---:|---:|
|
||
| **V1** | `movem.l` blit from a word-expanded RAM frame (96KB read + 96KB write) | **446,286** | **53.6%** |
|
||
| V2 | naive `move.b`/`move.w` per pixel from a byte source | 1,284,174 | 154.1% |
|
||
| **V3** | write-only floor — registers preloaded, no source read at all | **225,789** | **27.1%** |
|
||
| **V4** | the same 96KB of writes issued in **4x4 block order** | **637,971** | **76.6%** |
|
||
|
||
The 12fps budget is 833,333 cycles (10.0 MHz confirmed from `x68k.cpp:1133`,
|
||
`40_MHz_XTAL / 4`).
|
||
|
||
### 24.1 The numbers are cross-checked against hand-derived cycle counts
|
||
Every variant was predicted from the MC68000 timing tables *before* the run
|
||
(`MOVEM.L` M->R `(An)+` = 12+8n, `(d16,An)` = 16+8n; R->M `(An)` = 8+8n,
|
||
`(d16,An)` = 12+8n) and then measured:
|
||
|
||
| | predicted | measured | error |
|
||
|---|---:|---:|---:|
|
||
| V1 | 447,744 | 446,286 | 0.33% |
|
||
| V2 | 1,284,096 | 1,284,174 | 0.006% |
|
||
| V3 | 225,792 | 225,789 | 0.001% |
|
||
| V4 | 640,704 | 637,971 | 0.43% |
|
||
|
||
This agreement is the point. A MAME timing number on its own would be worth
|
||
little given how many false-good results this project has produced (FINDINGS 4);
|
||
two independent derivations landing within half a percent is worth something.
|
||
The residual error is the frame-granularity of the measurement — Lua gets no
|
||
cycle counter (`luaengine.cpp` exposes `machine.time` and nothing from
|
||
`device_execute_interface`), so timing resolution is one video frame, 18.03 ms.
|
||
|
||
### 24.2 SCOPE: these are instruction cycles, and therefore a LOWER BOUND
|
||
MAME's `gvram_w`/`gvram_r` (`x68k_crtc.cpp:501,595`) contain **no timing at
|
||
all** — no wait states, no `adjust_icount`. GVRAM in MAME is as fast as main
|
||
RAM. Real X68000 GVRAM stalls the CPU on access, so every figure above is a
|
||
floor, not a prediction. **Do not quote these as hardware numbers.** Interrupts
|
||
were masked (`SR = $2700`) so the IPL's timer and VBL handlers could not steal
|
||
cycles into the measurement; a real player will take interrupts on top.
|
||
|
||
### 24.3 The 38% estimate is dead — a full-frame blit is 53.6%
|
||
The realistic "decode into a RAM frame, then blit it" design costs **53.6% of
|
||
the frame budget before decoding a single block**, and that is the zero-wait-
|
||
state floor. The estimate the CPU budget has been carrying since session 1 was
|
||
38%. It was optimistic by 41%.
|
||
|
||
The cause is visible in the V1/V3 gap: **reading the source frame is exactly
|
||
half the total cost** (221,952 of 446,286 cycles). The 68000 pays 8 cycles per
|
||
longword read and 8 per longword written, and in 256-colour mode a pixel
|
||
occupies a whole word of address space, so a frame is 96KB of traffic in each
|
||
direction rather than 48KB.
|
||
|
||
### 24.4 The high byte of every GVRAM write is discarded — confirmed from source
|
||
`gvram_w` case `0x0100` writes `data & 0x00ff` with `mem_mask 0x00ff`. So in
|
||
256-colour mode the CPU cannot pack two pixels into one word, and the odd bytes
|
||
of a word-expanded source frame never need clearing — V1 exploits this by
|
||
leaving them uninitialised. This is why 96KB, not 48KB, is the irreducible
|
||
write traffic.
|
||
|
||
### 24.5 The architecture question, and where it turns over
|
||
> **Superseded by FINDINGS 28.1/28.2 (session 7).** The two-path plan below is
|
||
> incoherent — the compose path needs a RAM reference the direct path never
|
||
> writes — and its two costs are both *copies*, so they were never comparable to
|
||
> a decode. The "76.6% x non-SKIP fraction" model is also 2.03x optimistic:
|
||
> the four block modes cost 300/448/400 cycles, not one figure. One path ships.
|
||
|
||
V4 prices the access pattern a decoder that writes codewords **straight into
|
||
GVRAM** actually has: 4 rows of 8 bytes at a 1024-byte stride per 4x4 block. The
|
||
same 96KB of writes costs **76.6%** in block order versus 53.6% row-linear — the
|
||
stride destroys the `movem.l` burst, 208 cycles per block against a theoretical
|
||
best of ~150.
|
||
|
||
But a decoder never writes every block: SKIP blocks cost **nothing at all**, and
|
||
the previous frame is already sitting in GVRAM, so **no RAM reference frame is
|
||
needed for SKIP to work**. So the two designs scale differently:
|
||
|
||
- **compose-in-RAM then blit** — flat 53.6%, independent of how much changed
|
||
- **decode-direct-to-GVRAM** — 76.6% x (fraction of non-SKIP blocks)
|
||
|
||
**They cross at 70% of blocks changed.** Below that, writing straight into GVRAM
|
||
wins, and it also drops the 96KB RAM reference frame entirely. Above it, the
|
||
flat blit wins.
|
||
|
||
**This makes the non-SKIP block fraction the single most important unmeasured
|
||
number in the project.** It is already computable from the encoder — it is a
|
||
by-product of the mode decision in `vq_hybrid.py` — and it has never been
|
||
reported. Measure it before writing any decoder inner loop, because it selects
|
||
which inner loop to write.
|
||
|
||
### 24.6 The frame the 68000 drew is pixel-exact
|
||
V1's output was snapshotted and passes `verify_frame256.py` unchanged: `256x512
|
||
native, double-scan exact, active 256x192 pixel-exact, letterbox true black`,
|
||
40.81 dB. So 68000 code drives the mode of FINDINGS 23 correctly, and 23.5 is
|
||
now closed.
|
||
|
||
---
|
||
|
||
## 25. The sustained action sequence, found and measured (session 5)
|
||
|
||
STATUS has carried "a *sustained* action sequence is the one thing that could
|
||
still break the bitrate" as the open risk since session 2. Every clip measured
|
||
before this was 1.2-1.7 s. This section closes it: **it does break the profiles,
|
||
though not the bus.**
|
||
|
||
### 25.1 The two largest streams on the disc are not game footage
|
||
A survey that sorts 224 streams by size and encodes the biggest would have
|
||
measured **live action**:
|
||
|
||
| stream | size | what it actually is |
|
||
|---|---:|---|
|
||
| 00216 | 3777 MB | the feature with a **burned-in picture-in-picture commentary** |
|
||
| 00215 | 3475 MB | the commentary itself, full-screen live action |
|
||
| **00223** | **1802 MB** | **clean animation, 9.4 min — the one to use** |
|
||
|
||
The PiP in 00216 is burned into video stream 0, not a selectable secondary
|
||
stream, so there is no ffmpeg flag that recovers a clean frame from it. This
|
||
extends FINDINGS 13's menu-vs-content warning: the classification needed is
|
||
**content / menu / bonus**, and bonus material is the one that looks most like
|
||
content by every cheap metric (size, duration, bitrate).
|
||
|
||
### 25.2 Picking the worst window by measurement, not by eye
|
||
`tools/analysis/07_motion_survey.py` scans a whole stream at 96x72 and reports
|
||
the highest-mean sliding window of inter-frame absolute difference. On 00223:
|
||
|
||
```
|
||
6793 frames @12fps = 566.1s
|
||
motion energy mean 9.40 median 5.60 p90 21.70 max 112.39
|
||
hottest sustained 10s window: t = 539.4s (2.01x stream mean)
|
||
quietest 10s window: t = 144.2s (0.19x stream mean)
|
||
```
|
||
|
||
The 10.6x spread between the quietest and hottest sustained windows is the whole
|
||
argument for not sampling clips by hand. `t = 539.4s` is the Singe endgame.
|
||
|
||
### 25.3 Both profiles overshoot on that window — rate control is now required
|
||
Encoding those 120 frames at the shipping profiles, with the fixed `lam` the CLI
|
||
currently uses:
|
||
|
||
| profile | target | measured | overshoot | PSNR | palette ceiling |
|
||
|---|---:|---:|---:|---:|---:|
|
||
| `sasi` | 110 KB/s | **129.6 KB/s** | **+18%** | 27.82 dB | 31.33 dB |
|
||
| `scsi` | 280 KB/s | **373.8 KB/s** | **+34%** | 30.81 dB | 31.33 dB |
|
||
| *(00020 baseline, `sasi`)* | 110 KB/s | 108.0 KB/s | -2% | 36.94 dB | 39.90 dB |
|
||
|
||
**This reclassifies rate control from insurance to a requirement.** STATUS has
|
||
had "wire rate control into `encode.py`" at priority 3-4 since session 2 with the
|
||
note "no longer a blocker (FINDINGS 21)". That was true of the clips measured
|
||
then. It is not true of this one. `ratectl.encode_rate_controlled()` already
|
||
exists and builds a per-frame lam ladder; it has simply never been hooked up.
|
||
|
||
Note what did **not** break: 373.8 + 7.8 = 381.6 KB/s is still under the 488 KB/s
|
||
working figure, so FINDINGS 21's ring-buffer conclusion survives — but at 78% of
|
||
the pipe sustained over ten seconds rather than the comfortable margin implied by
|
||
1.7 s clips.
|
||
|
||
### 25.4 The palette ceiling is content-dependent, and on hard content it binds
|
||
The 256-colour scene palette costs **31.33 dB** on this window against **39.90 dB**
|
||
on 00020 — 8.6 dB worse. Fire, lava and smoke gradients are exactly what a
|
||
256-entry mediancut palette handles worst.
|
||
|
||
This inverts an assumption the project has been carrying. FINDINGS 23.3 put the
|
||
X68000 display ceiling at 40.81 dB and treated it as comfortably clear of the
|
||
codec's own error. On this content the **scene palette (31.33 dB), not the
|
||
display hardware (40.81 dB), is the binding constraint** — and `scsi` is already
|
||
within 0.51 dB of it. Spending bits to close that last half-dB is spending them
|
||
against a ceiling that is not the display's.
|
||
|
||
### 25.5 `scsi` collapses to RAW under stress
|
||
Mode distribution on this window is qualitatively different from anything
|
||
measured before:
|
||
|
||
| profile | SKIP | V1 | V4 | RAW |
|
||
|---|---:|---:|---:|---:|
|
||
| `sasi` (lam=60) | 45.6% | 16.3% | 24.2% | 13.9% |
|
||
| `scsi` (lam=10) | 26.2% | 5.5% | 7.1% | **61.2%** |
|
||
| *00020, `sasi`* | 46.9% | 24.1% | 17.8% | 11.2% |
|
||
|
||
At `lam=10` the rate-distortion decision finds literal pixels cheaper than any
|
||
codeword for 61% of blocks — the codebooks are simply not describing this
|
||
content. That is the mechanism behind the +34% overshoot in 25.3, and it is a
|
||
rate-control problem, not a codec-structure problem: the RD decision is behaving
|
||
correctly for the lam it was given.
|
||
|
||
### 25.6 The decoder needs BOTH display paths, chosen per frame
|
||
> **Superseded by FINDINGS 28.1 (session 7).** Mixing the paths displays stale
|
||
> pixels on 70 of these 120 frames. The "median 37.0%, capped at 53.6%" below is
|
||
> the cost of an incorrect player; every coherent version is dearer, and plain
|
||
> direct-to-GVRAM is the cheapest of them.
|
||
|
||
Applying FINDINGS 24.5's crossover to the real per-frame distribution:
|
||
|
||
| | median non-SKIP | p90 | frames over the 70% crossover |
|
||
|---|---:|---:|---:|
|
||
| `sasi`, Singe window | 48.4% | 82.8% | 36 / 120 (30%) |
|
||
| `scsi`, Singe window | 70.8% | 92.4% | 64 / 120 (53%) |
|
||
| `sasi`, 00020 | 54.0% | 88.8% | 3 / 14 (21%) |
|
||
|
||
Neither path wins outright: **30-53% of frames want the flat blit and the rest
|
||
want direct-to-GVRAM.** A player that implements both and picks per frame — the
|
||
mode headers are parsed before any pixel is written, so the count is free — pays
|
||
a median of **37.0%** of the frame budget and is capped at **53.6%**. A player
|
||
that implements only direct-to-GVRAM pays up to 76.6% and would miss frames on
|
||
the scene cuts.
|
||
|
||
So the answer to 24.5 is "both", and the selection is a one-line comparison
|
||
against a block count the decoder already has in hand.
|
||
|
||
### 25.7 What this does not measure
|
||
One 10 s window of one stream, at fixed lam, with `_paint` still a Python loop.
|
||
The full-disc survey is still not done, and the numbers above are the *worst*
|
||
window rather than a distribution over content. What has changed is that the
|
||
worst case is now a measurement rather than a worry.
|
||
|
||
---
|
||
|
||
## 26. Rate control is unsound as written — found before wiring it up (session 5)
|
||
|
||
FINDINGS 25.3 promoted rate control from insurance to a requirement. Reading
|
||
`ratectl.py` before wiring it into `encode.py` turned up a correctness bug that
|
||
would have produced exactly the kind of plausible-looking wrong result this
|
||
project keeps catching (FINDINGS 4, 9, 14, 18).
|
||
|
||
### 26.1 The lam ladder desynchronises the encoder from the decoder
|
||
`H.encode()` is **temporally recursive**: SKIP blocks are copied from the
|
||
previous *reconstruction*, and `prev = out` closes the loop
|
||
(`vq_hybrid.py:84-109`). A frame's output therefore depends on every frame
|
||
before it in that same run.
|
||
|
||
`encode_rate_controlled()` runs `H.encode()` once per lam over the **whole
|
||
sequence**, building a ladder of independent temporal chains, then picks each
|
||
frame from whichever rung fits the budget. When frame *f* comes from rung *i*
|
||
and frame *f-1* was emitted from rung *j != i*, the SKIP blocks in *f* reference
|
||
a reconstruction **the decoder never saw**.
|
||
|
||
Measured on the Singe window (`tools/analysis/09_ratectl_drift.py`, 120 frames,
|
||
5 rungs, target 110 KB/s):
|
||
|
||
| | |
|
||
|---|---|
|
||
| rung switches | **67** over 120 frames |
|
||
| frames whose emitted output differs from what the encoder recorded | **111 / 120** |
|
||
| worst frame | **21,339 px = 43.4% of the frame** |
|
||
| encoder-vs-decoder agreement, worst frame | 27.1 dB |
|
||
| reported PSNR overstatement | **0.36 dB** |
|
||
|
||
The 0.36 dB is the least interesting number here. The encoder is reporting
|
||
quality for a reconstruction that will never exist, and 43% of a frame differing
|
||
is a visible artefact whatever the mean says.
|
||
|
||
**The fix is structural, not a tuning change:** `H.encode()` must become
|
||
frame-drivable — take `prev` and one lam, return one frame — so rate control can
|
||
feed back the frame it actually emitted. The current whole-sequence signature is
|
||
what makes the ladder tempting in the first place.
|
||
|
||
### 26.2 The ladder spans 250x past the shippable range
|
||
`lam_hi=2e5`, but FINDINGS 15 puts the quality cliff between lam=800 and
|
||
lam=2000 and says do not ship past lam~800. Every rung above ~800 is
|
||
unshippable, so a frame that only fits at lam=9457 has not been rate-controlled,
|
||
it has been destroyed. Cap `lam_hi` at 800 and let a frame that cannot fit
|
||
overrun the bucket — a visible overrun is a better failure than silent garbage.
|
||
|
||
### 26.3 The ladder is far too coarse where it matters
|
||
With `steps=5` the geomspace lands on 1 / 21 / 447 / 9457 / 200000, and **only
|
||
two rungs were ever chosen**. The budget is 8,721 B/frame; the two straddling
|
||
rungs deliver 23,183 B (lam=21) and 3,071 B (lam=447) — a **7.5x** gap across
|
||
the operating point. Rate control cannot land near a target it has to jump over.
|
||
|
||
The module docstring already describes the right approach — *"per frame we
|
||
binary-search lam to land inside a byte budget"* — but the implementation is a
|
||
fixed precomputed ladder. Doc and code disagree; the doc is correct.
|
||
|
||
### 26.4 What does work
|
||
The leaky bucket lands the mean where it should: **109.1 KB/s against a 110
|
||
target**, with 32% of frames over the per-frame budget and banked by the bucket.
|
||
That mechanism is sound and worth keeping. It is the per-frame lam *selection*
|
||
underneath it that needs rebuilding, not the bucket.
|
||
|
||
### 26.5 Cost note before starting
|
||
Each rung is a full-sequence encode and `_paint` is still a Python per-block
|
||
loop, so a 5-rung run over 120 frames takes minutes. **Vectorise `_paint`
|
||
first** — it is already on the list for the full-disc survey and it makes the
|
||
rate-control work practical rather than merely faster.
|
||
|
||
## 27. Rate control, rebuilt and wired in (session 6)
|
||
|
||
FINDINGS 26 stopped the session-5 rate controller before it shipped: it picked
|
||
frames out of independently-encoded whole-sequence runs, so 111 of 120 frames
|
||
referenced reconstructions the decoder would never see. The fix was structural,
|
||
as 26.1 said it had to be. It is now wired into `encode.py` and **on by
|
||
default** for a profile.
|
||
|
||
### 27.1 The encoder is frame-drivable, and the drift is zero by construction
|
||
`vq_hybrid` now exposes one frame at a time — `frame_ctx(m, f, prev)` /
|
||
`decide(ctx, lam)` / `paint(m, ctx, mode)` — and `encode()` is a thin loop over
|
||
that API. Rate control drives the same three calls and feeds back **the frame it
|
||
actually emitted** as the next frame's `prev`. There is no ladder to pick from,
|
||
so the desync has no way to occur.
|
||
|
||
`tools/analysis/09_ratectl_drift.py`, unchanged in what it asserts:
|
||
|
||
| | session 5 | session 6 |
|
||
|---|---|---|
|
||
| frames whose emitted output differs from what the encoder recorded | 111 / 120 | **0 / 120** |
|
||
| worst frame | 21,339 px (43.4%) | **0 px** |
|
||
| reported PSNR overstatement | 0.36 dB | **0.00 dB** |
|
||
|
||
This is the harder case for that test on purpose: it runs with `lam_lo=1.0`, so
|
||
lam moves on 117 of 119 frame boundaries. Under the old ladder, 67 rung switches
|
||
were enough to corrupt 111 frames.
|
||
|
||
### 27.2 Both overshoots are closed, and they cost under 1 dB
|
||
The Singe window (FINDINGS 25.3), which is the worst sustained window on the
|
||
disc. Totals include the 7.8 KB/s ADPCM allowance:
|
||
|
||
| profile | target | fixed lam (session 5) | rate-controlled | quality cost |
|
||
|---|---|---|---|---|
|
||
| `sasi` | 110 KB/s | 137.4 KB/s (**+25%**) | **109.5 KB/s** | 27.82 → 27.22 dB (−0.60) |
|
||
| `scsi` | 280 KB/s | 381.6 KB/s (**+36%**) | **280.0 KB/s** | 30.81 → 29.90 dB (−0.91) |
|
||
|
||
Zero frames hit the lam=800 cliff at either profile, so nothing was destroyed to
|
||
get there (26.2's failure mode did not trigger). `sasi` needed lam to reach 183
|
||
at worst against a floor of 60; `scsi` reached 58.7 against 10. The controller
|
||
is working an order of magnitude below the cliff, which is where the search
|
||
range being capped at 800 rather than 2e5 stops mattering at all — and that is
|
||
the point: a range that never needs its top is a range you can trust.
|
||
|
||
`scsi` still sits **1.43 dB** from the scene palette ceiling of 31.33 dB
|
||
(FINDINGS 25.4), against 0.51 dB before. The ceiling, not the codec, is still
|
||
what bounds this content.
|
||
|
||
The percentages differ from 25.3's +18%/+34% because those compared video
|
||
payload against the total target; the table above compares like with like
|
||
(total against total). The payload figures are unchanged: 129.6 and 373.8 KB/s.
|
||
|
||
### 27.3 Rate control makes the display path cheaper, not dearer
|
||
The decoder-architecture numbers of FINDINGS 25.6 were measured on the
|
||
fixed-lam encoder. Re-measured under rate control, on the same window, with the
|
||
player picking the cheaper of compose-then-blit and direct-to-GVRAM per frame:
|
||
|
||
| profile | median display cost | frames above the 70% crossover |
|
||
|---|---|---|
|
||
| `sasi` fixed → RC | 37.0% → **36.6%** | 30.0% → 26.7% |
|
||
| `scsi` fixed → RC | 53.6% → **47.1%** | 53.3% → 35.8% |
|
||
|
||
Raising lam moves blocks to SKIP and V1, which is fewer blocks to write. The
|
||
"implement both paths, pick per frame" conclusion is unaffected and the cap is
|
||
still 53.6%.
|
||
|
||
### 27.4 The quality floor barely matters; the prefill matters, wrongly
|
||
Two knobs were measured rather than guessed.
|
||
|
||
**`--rc-floor`** decides whether a quiet frame may spend more than the fixed-lam
|
||
profile would. On the Singe window it is worth nothing — 109.5 vs 110.0 KB/s and
|
||
**0.00 dB** — because no frame on that window is quiet enough for the bucket to
|
||
saturate. The default is `profile` (never spend more than session 5 would), so
|
||
rate control cannot regress content that already fits.
|
||
|
||
**`--prefill`** models how full the player's buffer is at scene start. It is
|
||
tempting and it is a trap, so it defaults to 0:
|
||
|
||
| clip | prefill 0.0 | 0.5 | 1.0 | target |
|
||
|---|---|---|---|---|
|
||
| Singe, 120 fr, `sasi` | 109.5 | 112.9 | **116.3** | 110 |
|
||
| Singe, 120 fr, `scsi` | 280.0 | 289.1 | **298.2** | 280 |
|
||
| 00020, 14 fr, `sasi` | 92.0 | **115.8** | **115.8** | 110 |
|
||
| 00020, 14 fr, `scsi` | 224.8 | **255.9** | **255.9** | 280 |
|
||
|
||
(`scsi` on 00020 is the one cell where prefill looks harmless: the clip fits
|
||
under 280 either way. That is the content being easy, not the knob being safe.)
|
||
|
||
Prefill buys a permission to overshoot of exactly `bucket / nframes`. At 8
|
||
frames of bucket over 120 frames that is 6.2% — measured — and on a 14-frame
|
||
clip the bucket is larger than the clip, so rate control switches itself off and
|
||
reproduces fixed-lam exactly (lam never leaves its floor: min = median = max =
|
||
60). **A prefill that makes a target look met has disabled the controller.**
|
||
|
||
### 27.5 The 00020 undershoot is a clip-length artefact, not a bug
|
||
At prefill 0 the 14-frame 00020 clip lands at 92.0 KB/s against a 110 ceiling —
|
||
0.66 dB given away for nothing. That is the leaky bucket's startup transient:
|
||
the first `bucket_frames` frames cannot draw on a bank they have not accumulated.
|
||
It is bounded by `bucket / nframes`, so it is 6% on a 10-second window and 20%
|
||
on a 1.2-second one.
|
||
|
||
The lesson is the one FINDINGS 25.3 already taught in a different costume: **a
|
||
1.2-second clip cannot be used to judge rate control.** Real scenes are tens of
|
||
seconds. Do not tune the bucket against 00020.
|
||
|
||
Worth recording separately: fixed-lam `sasi` on 00020 delivers 115.8 KB/s — the
|
||
supposedly easy clip was **already 5% over its target**, which nothing had
|
||
noticed because the profile table quotes its PSNR and not its bitrate.
|
||
|
||
### 27.6 FINDINGS 26.5's cost premise was wrong in both halves
|
||
26.5 said a rate-control experiment was minutes because `_paint` is a Python
|
||
per-block loop, and told the next session to vectorise it first. Vectorising it
|
||
was correct and it is **17.1x faster**, but it was never the bottleneck, and the
|
||
ladder was never minutes. Measured per frame, 256x192:
|
||
|
||
| | ms |
|
||
|---|---|
|
||
| `VQ.assign` x2 — codeword search | **22.83** |
|
||
| SKIP error against `prev` | 1.40 |
|
||
| `decide` — argmin at one lam | 0.06 |
|
||
| `paint`, vectorised | 0.29 |
|
||
| `paint`, old per-block loop | 4.93 |
|
||
|
||
`_paint` was 14% of a frame. A 5-rung ladder over 120 frames was ~18 s of
|
||
encoding, not minutes — the "few minutes" in the drift test's docstring was
|
||
`H.build`'s k-means (51 s), which no amount of vectorising `_paint` would have
|
||
touched.
|
||
|
||
What actually makes per-frame rate control affordable is that `VQ.assign`'s
|
||
output depends on **neither `lam` nor `prev`**, so it is computed once per frame
|
||
and a lam search only re-runs the 0.06 ms argmin:
|
||
|
||
| | |
|
||
|---|---|
|
||
| 12-step per-frame lam search, 120 frames, symbols cached | **0.31 s** |
|
||
| the same search by re-running whole-sequence encodes | 49.10 s |
|
||
|
||
That is a 158x difference, and it is the reason the controller can afford a real
|
||
bisection instead of a 5-rung ladder — which was the actual defect in 26.3.
|
||
|
||
The cache holds **one frame**. At ~133 KB of intermediates per frame, caching
|
||
the sequence would cost 900 MB on a 9.4-minute stream to save nothing: every
|
||
caller works a frame at a time.
|
||
|
||
---
|
||
|
||
## 28. The 68000 decoder exists, is pixel-exact, and does not fit (session 7)
|
||
|
||
`src/player/decode.s` parses DLX1 and draws frames on the emulated X68000. It is
|
||
**pixel-exact across a 120-frame sequential run** of the worst sustained window
|
||
on the disc (`tools/bench/verify_decode.py`), exercising all four block modes
|
||
and the full temporal recursion — the last frame is only right if every frame
|
||
before it was.
|
||
|
||
It is also **too slow**. On that window, at the shipping `sasi` profile:
|
||
|
||
| | non-SKIP blocks | measured cost |
|
||
|---|---:|---:|
|
||
| cheapest frame | 15.4% | 31.5% of a 12fps frame |
|
||
| median frame | 47.8% | 73.8% |
|
||
| p90 frame | 82.5% | **116.4%** |
|
||
| worst frame | 100.0% | **135.8%** |
|
||
| mean over the window | 47.8% | **81.7%** |
|
||
|
||
**31% of frames miss the 833,333-cycle budget**, and like every figure since
|
||
FINDINGS 24 these are instruction cycles against zero-wait-state memory, so
|
||
they are a floor. This is the first time CPU, not disk, is the binding
|
||
constraint — FINDINGS 21 retired the bandwidth worry, and this replaces it.
|
||
|
||
### 28.1 The dual-path plan of 24.5/25.6 was incoherent, and is withdrawn
|
||
FINDINGS 24.5 specified two display paths chosen per frame on the non-SKIP
|
||
count, and 25.6 costed the mix at "median 37.0%, capped at 53.6%". Two of its
|
||
premises cannot both hold:
|
||
|
||
- compose-in-RAM-then-blit exists to make the blit **row-linear**, so it must
|
||
assemble a **full** frame in RAM. The pixels it does not decode this frame —
|
||
the SKIP blocks — can only come from a RAM copy of the previous
|
||
reconstruction.
|
||
- decode-direct-to-GVRAM's stated advantage is that **"no RAM reference frame
|
||
is needed"**, because the previous frame is already in GVRAM.
|
||
|
||
So every direct frame silently invalidates the reference the next compose frame
|
||
reads. Simulated on the Singe window at the crossover the plan specifies
|
||
(`tools/analysis/10_pathmix_drift.py`): **70 of 120 frames display pixels no
|
||
correct player would display**, first at frame 2, worst frame 18.8% of the
|
||
screen. This is FINDINGS 26 in different clothing — two code paths disagreeing
|
||
about what "the previous frame" means — and it is the **sixth** false premise
|
||
this project has caught before it shipped.
|
||
|
||
Every coherent repair is worse than not mixing at all:
|
||
|
||
| strategy | median | p90 | max | correct |
|
||
|---|---:|---:|---:|---|
|
||
| mix per frame, as specified | 36.6% | 53.6% | 53.6% | **no** |
|
||
| mix, direct also writes the RAM reference | 53.6% | 68.4% | 81.4% | yes |
|
||
| mix, re-read GVRAM into RAM on each switch | 36.6% | 107.2% | 107.2% | yes, 13 frames miss |
|
||
| compose only | 53.6% | 53.6% | 53.6% | yes |
|
||
| **direct only** | **36.6%** | 62.5% | 76.6% | yes |
|
||
|
||
(Costs in that table are 24.5's own model, for like-for-like comparison; 28.2
|
||
replaces the model itself.)
|
||
|
||
**24.5 also compared the wrong two things.** Its 53.6% and 76.6% are both
|
||
*copies* measured in `blit.s` — neither includes decoding. A real compose path
|
||
costs decode-into-RAM **plus** the 53.6% blit, so it is strictly dearer than
|
||
decoding straight into GVRAM, whatever the block mix. There was never a
|
||
crossover to find.
|
||
|
||
**The decoder therefore implements one path, direct-to-GVRAM**, and drops the
|
||
96 KB RAM reference frame entirely.
|
||
|
||
### 28.2 The four block modes do not cost the same, and V4 is the expensive one
|
||
24.5's model — "76.6% of a frame x the non-SKIP fraction" — prices every
|
||
non-SKIP block as one `movem.l` burst. Measured separately, with synthetic
|
||
single-mode frames (`tools/bench/prep_dlx.py`):
|
||
|
||
| mode | cycles/block | vs the 24.5 model (207.8) |
|
||
|---|---:|---:|
|
||
| SKIP, in an all-SKIP header byte | 13.3 | model says 0 |
|
||
| SKIP, inside a mixed byte | ~45 | model says 0 |
|
||
| V1 (one 4x4 codeword) | **299.9** | 1.44x |
|
||
| V4 (four 2x2 codewords) | **448.2** | 2.16x |
|
||
| RAW (16 literal indices) | **400.4** | 1.93x |
|
||
|
||
Applied to the real per-frame histograms (`tools/analysis/11_cpu_budget.py`),
|
||
the model reproduces all four frames timed on the 68000 to within **1
|
||
percentage point**, and shows 24.5 to be **2.03x optimistic at the median**.
|
||
|
||
Where the cycles actually go over the window:
|
||
|
||
| mode | % of blocks | % of cycles |
|
||
|---|---:|---:|
|
||
| SKIP | 46.4% | 9.2% |
|
||
| V1 | 19.8% | 26.1% |
|
||
| V4 | **25.2%** | **49.7%** |
|
||
| RAW | 8.5% | 15.0% |
|
||
|
||
**V4 is a quarter of the blocks and half the cycles.** It costs 1.49x a V1 block
|
||
while the mode decision in `vq_hybrid.py` charges it only its 4x payload bytes.
|
||
The lagrangian trades distortion against *bytes*; on this machine it now has to
|
||
trade distortion against *cycles* as well.
|
||
|
||
### 28.3 The container is big-endian but not aligned, and that is an address error
|
||
The DLX1 header docstring says every multi-byte field is big-endian "so the
|
||
68000 reads them with a plain `move`". Alignment is the other half of that
|
||
sentence and the container does not have it: frame records are
|
||
`[u32 length][768-byte mode header][payload]` laid end to end with arbitrary
|
||
payload lengths, so record boundaries land on odd addresses.
|
||
|
||
`move.l (a0)+,d0` at an odd address is an **address error** on a 68000 — not a
|
||
slow read. The first run decoded frame 0 perfectly, consumed exactly its 8,715
|
||
payload bytes, then read frame 1's length at `$03220F` and vectored into the IPL
|
||
at `$FF059A`, where it sat for 59 emulated seconds looking like an infinite
|
||
loop. It was found by dumping PC and the address registers, not by reading the
|
||
source: the code was correct, the data layout was not.
|
||
|
||
The decoder now rounds each record start up to 4. **The container should carry
|
||
the padding itself** so a streaming player can DMA records into place: measured
|
||
cost on this window is **199 bytes over 120 frames — 1.66 B/frame, 20 B/s**
|
||
against a 110 KB/s budget. Until `encode.py` does it, `prep_dlx.py` realigns at
|
||
load time.
|
||
|
||
### 28.4 The measurements agree with hand-derived MC68000 timings
|
||
As in FINDINGS 24, each figure was derived from the instruction timing tables
|
||
before being believed. A V1 block, summing dispatch, index decode, the indexed
|
||
`movem.l` load and four `movem.l` stores, plus its quarter share of the header
|
||
loop: **298.5 cycles derived against 299.9 measured — 0.5%.** RAW derives to
|
||
396 against 400.4 measured (1%). V4 derives to 415 against 448 (7%, the gap
|
||
being in the indexed two-register `movem.l`, the mode this decoder uses most
|
||
heavily). So these are 68000 cycles, not a MAME artefact.
|
||
|
||
### 28.5 A full frame does not fit at 12fps in ANY mode
|
||
An all-V1 frame — the cheapest possible way to redraw all 3,072 blocks — costs
|
||
**921,187 cycles, 110.5% of the budget**. All-V4 is 165.2% and all-RAW 147.6%.
|
||
|
||
So the ceiling is structural, not a tuning problem: **at 12fps on a 10MHz 68000
|
||
no more than ~88% of the screen can change in one frame**, however cheaply it is
|
||
coded. Scene cuts change 100% of it. Either a cut gets one late frame (the
|
||
outgoing content is unrelated, so this may be free to the eye), or cuts have to
|
||
be spread across two frame times, or the framerate has to come down — at 10fps
|
||
the budget is 1,000,000 cycles and an all-V1 frame fits.
|
||
|
||
### 28.6 What this does not measure
|
||
One 10 s window of one stream at one profile, and MAME still models no GVRAM
|
||
wait states. The `scsi` profile will be worse: FINDINGS 25.5 has it collapsing
|
||
to RAW under stress, and RAW is 1.93x the old model's block. Nothing here has
|
||
been run on `00020` or on quiet content, where the median frame is far cheaper.
|