User clarified the bandwidth figure is 4 Mbps (488 KB/s), not 4 MB/s -- ~8x
tighter than the previous commit reasoned against. Two consequences, plus a
correction to session 1.
1. The scsi profile committed in f0f2f80 DOES NOT FIT. Its mean is a
comfortable 52% of the pipe but it PEAKS at 96.4% (470.8 KB/s on scene
00020), and a frame that arrives late is a dropped frame, not a slow one.
Peak/mean is 1.4-1.9x even on 1.2-1.7s clips. Sizing a real-time stream on
the mean was the error. Flagged in STATUS rather than silently retuned,
because the fix is rate control, not a lower lam.
This promotes ratectl.py -- written in session 2, never wired into
encode.py -- from a loose end to the highest-value work in the repo. It is
worth a full step on the quality ladder (lam=25 -> lam=10, +0.7/+1.2 dB)
because it allows sizing for the mean instead of the peak.
2. Pixel-exact is off the table at this bandwidth: lam=0 needs 92-97% of the
pipe. The previous commit's "if SCSI sustains >=800 KB/s, ship transparent"
conclusion only applies at roughly double the user's figure.
3. FINDINGS 5 said that because transfers are DMA, streaming "costs essentially
no CPU" and the 68000 is "nearly idle". That is wrong. The HD63450 steals
~8 clocks per 16-bit word: 10-20% of the machine at the rates the profiles
now use, on top of a 38% full-frame blit. Bandwidth and CPU are one budget.
Adds tools/encoder/profile_gen.py, which derives lam FROM a bandwidth figure
(accounting for audio, peak/mean and DMA steal) instead of reading it off the
knee of the RD curve, and docs/BENCHMARK.md covering how to actually measure
the subsystem -- including why MAME cannot answer the bandwidth question and
would be the same class of error as the FINDINGS 4 traps.
The 4 Mbps figure is user-supplied and its provenance is not recorded; every
profile now hangs off it.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
514 lines
24 KiB
Markdown
514 lines
24 KiB
Markdown
# Findings — session 1 (2026-08-23)
|
|
|
|
All numbers here are MEASURED unless marked ESTIMATE or FOLKLORE.
|
|
|
|
---
|
|
|
|
## 1. Source material
|
|
|
|
`DRAGONS_LAIR.iso` — 16 GB, UDF 2.x, **decrypted** (no AACS dir).
|
|
Loop-mounted read-only at `/media/reala-misaki/BDROM` via `udisksctl loop-setup -r -f`.
|
|
(7-Zip cannot read UDF 2.x; use the loop mount.)
|
|
|
|
- **224 `.m2ts` streams**, 1920x1080, **MPEG-2, progressive, 23.976 fps**
|
|
- Size histogram: 47 <5MB, 138 5-50MB, 22 50-150MB, 14 150-400MB, 3 >400MB
|
|
- The 185 sub-50MB streams are the **arcade branching scenes already split into
|
|
individual clips** — we get scene boundaries for free.
|
|
- Big streams are full-feature playthroughs: 00215 (1376s), 00216 (1151s), 00223 (566s)
|
|
- Typical scene clip ~60s (00203/00205/00199), some ~100s (00164/00212)
|
|
|
|
**Gotcha:** clip durations vary wildly. Always read `format=duration` and seek
|
|
relative to it. Seeking to a fixed offset silently yields 0 frames on short clips.
|
|
|
|
---
|
|
|
|
## 2. GVRAM layout [verified — see HARDWARE.md for source]
|
|
|
|
**One 16-bit word per pixel position in EVERY color mode.** Bit depth does not
|
|
change VRAM bandwidth; it only subdivides the word.
|
|
|
|
`addr = page_base + y*1024 + x*2` — adjacent pixels are 2 bytes apart in all modes.
|
|
|
|
Consequence: low bit depth buys **no speed**. 16-color mode is strictly worse than
|
|
256-color (same bus traffic, 1/16 the palette). Page-alias writes are hardware
|
|
auto-masked, so 16-color needs no software read-modify-write — but it's still
|
|
one word-access per pixel.
|
|
|
|
**Chosen: 256 colors, 256x192 active area.**
|
|
In 256-color mode P0=low byte, P1=high byte of each word. Sacrificing page 1 as a
|
|
double-buffer lets a `move.l` cover two pixel positions, enabling `movem.l` bursts
|
|
(12 regs = 48 bytes = 24 pixels). Identical blit cost to 65536-color mode but
|
|
**half the on-disk data**.
|
|
|
|
---
|
|
|
|
## 3. Content measurements (8 scenes sampled, 5s each at 40% into each clip)
|
|
|
|
| metric | mean | p90 |
|
|
|---|---|---|
|
|
| pixels changed / frame | 20.1% | 30.2% |
|
|
| **blit cost** | **~64k cycles** | **~97k cycles** |
|
|
| naive delta+RLE frame size | 15.5 KB | 19.6 KB |
|
|
|
|
Budget is **833,333 cycles/frame** @ 12fps on a 10MHz 68000.
|
|
|
|
### => THE CPU IS NOT THE BOTTLENECK. I/O IS.
|
|
Blit uses **under 8%** of budget. The naive row-span+RLE codec achieves only
|
|
**3.2:1**, giving **365 KB/s / 470 MB** at 24fps (~183 KB/s / 235 MB at 12fps).
|
|
|
|
Per-scene variance is extreme: static dialogue ~30 KB/s, action ~700 KB/s.
|
|
Any codec needs a hard bitrate ceiling, not just a good average.
|
|
|
|
### "Shot on twos" — ASSUMPTION FAILED
|
|
Dedupe found **zero** duplicate frames across all 8 scenes (`uniq=120/120`,
|
|
24.0 fps effective). This Blu-ray is a restoration where every frame is unique.
|
|
We do NOT get halved data for free. **Decimation to 12fps must be explicit.**
|
|
|
|
A weak alternation signature does exist (even-index pairs 40.7% vs odd 27.5%,
|
|
ratio 1.5x, with occasional true-duplicate pairs at 0.03-0.19%), but it is
|
|
irregular — Bluth mixed ones and twos; action is animated on ones.
|
|
|
|
---
|
|
|
|
## 4. MEASUREMENT TRAPS — read before trusting any pipeline number
|
|
|
|
Three separate false results were produced and caught this session. All three
|
|
looked plausible. Guard against them:
|
|
|
|
1. **Per-frame Floyd-Steinberg dithering destroys temporal coherence.**
|
|
Error diffusion is chaotic: a +/-1 input change cascades across the row and
|
|
produces a completely different index pattern. First run reported 31.5% pixels
|
|
changed with near-zero variance (median 31.6, p90 32.3, max 32.7) while source
|
|
mean-abs-diff was 0.09 — i.e. visually identical frames. That flat variance is
|
|
the tell: **real animation has scene-dependent variance; noise does not.**
|
|
Use no dithering (cel art is flat) or ordered/Bayer (spatially fixed, temporally stable).
|
|
|
|
2. **Temporal denoise smears motion.** `hqdn3d=4:3:6:4` — the `6:4` are temporal
|
|
params. It flattened real motion, which then measured as "no motion" and
|
|
produced an absurd 0.8 fps / 4 MB result. **Use spatial-only: `hqdn3d=4:3:0:0`.**
|
|
|
|
3. **Exact-match dedupe fails on a noisy source.** MPEG-2 grain means near-duplicate
|
|
frames differ by +/-1 and are never bit-exact. Use a threshold on
|
|
"% pixels differing by more than N levels", and pick the threshold from the
|
|
observed distribution, not a guess. A 2% threshold ate genuine animation when
|
|
mean consecutive change was only 0.9%.
|
|
|
|
**Sanity rule: if a result has suspiciously low variance, or is suspiciously
|
|
good, it is probably an artifact of the measurement, not a property of the content.**
|
|
|
|
Scripts kept in `tools/analysis/` — 01 and 02 are marked BROKEN deliberately as
|
|
regression references; 03 and 04 are the correct ones.
|
|
|
|
---
|
|
|
|
## 5. Storage interface — the SASI/SCSI split
|
|
|
|
[Yasuma, X68030 internal SCSI controller]
|
|
|
|
- Interface: **SCSI-1**, 50-pin, 5 MB/s bus spec
|
|
- Controller: **Fujitsu MB89352** SPC
|
|
- Transfer mode: **DMA** (via **HD63450** DMAC)
|
|
- Bus: X68000 original bus, **16-bit @ 10MHz**
|
|
|
|
**Even on the X68030, SCSI runs at 10MHz 16-bit DMA.** Storage bandwidth does
|
|
NOT scale with CPU — the controller sits on the original bus. HD63450's 12.5MHz
|
|
official ceiling is why the X68030 runs at 25MHz. An "HSCSI" TSR forces PIO/FIFO
|
|
transfer instead of DMA but was marginal even at 25MHz.
|
|
|
|
Because it's DMA, **streaming costs essentially no CPU** — this stacks with the
|
|
8% blit utilisation. The 68000 really is nearly idle.
|
|
|
|
### Model split — IMPORTANT
|
|
**The 10MHz models (original X68000, ACE, PRO, EXPERT) use SASI, not SCSI.**
|
|
Built-in SCSI starts at the X68000 **Super** (1990) and continues through XVI,
|
|
Compact, X68030. SCSI on earlier machines needs the **Sharp CZ-6BS1** board
|
|
in an I/O slot (MAME models this: `-exp1 cz6bs1`).
|
|
|
|
| target | bandwidth | naive codec (365 KB/s) | VQ codec (~30 KB/s) |
|
|
|---|---|---|---|
|
|
| SASI (stock ACE/EXPERT) | ~300-500 KB/s FOLKLORE | infeasible | comfortable |
|
|
| SCSI (Super+, or CZ-6BS1) | ~1 MB/s FOLKLORE | tight but viable | trivial |
|
|
|
|
Derived bounds (ESTIMATE): 16-bit @10MHz with 4-clock bus cycle = 5 MB/s absolute
|
|
ceiling; HD63450 single-address DMA ~8 clocks/word => ~2.5 MB/s practical ceiling,
|
|
before SCSI-1 async handshake and drive latency.
|
|
|
|
**No measured benchmark was obtained — see STATUS.md.** The ~300-500 KB/s and
|
|
~1 MB/s figures are folklore-grade; I could not find a primary measurement.
|
|
|
|
---
|
|
|
|
## 6. Codec decision: vector quantization (Cinepak-style)
|
|
|
|
Given ~8x CPU headroom and an I/O ceiling, spend CPU to buy bandwidth.
|
|
|
|
- Split frame into 4x4 blocks, encode each as a 1-byte index into a per-scene codebook
|
|
- Decode = 16-byte copy from a lookup table: nearly free
|
|
- A **full** frame = 256*192/16 = **3,072 bytes** — a hard 16:1 floor before delta
|
|
- Add block-level delta on top; action scenes ~2-3 KB/frame
|
|
- => roughly **30 KB/s, ~40 MB total**, with a *deterministic* bitrate ceiling
|
|
|
|
Divergence from the SNES project (below): use a **per-scene codebook with delta
|
|
updates**, not a per-frame rebuild. We trade adaptivity for bandwidth because we
|
|
have 2MB RAM to keep a codebook resident and CPU to spare.
|
|
|
|
**Risk not yet evaluated:** 4x4 VQ with a 256-entry codebook will visibly soften
|
|
detail. Bluth's fine ink linework is what suffers. Prototype and eyeball before committing.
|
|
|
|
---
|
|
|
|
## 7. Comparison: astrobleem/SNES-SuperDragonsLairArcade
|
|
|
|
Reached the **same core architecture independently** — "512 tiles per frame" is
|
|
vector quantization (8x8 codebook + tilemap). Good validation.
|
|
|
|
But: the SNES PPU has **no bitmap mode**, so tiles are forced on them by display
|
|
hardware. The X68000 has a real linear framebuffer, so VQ is a *compression
|
|
choice* we can tune or drop per-scene.
|
|
|
|
**MSU-1 is a bandwidth cheat we don't have.** It's a modern flash-cart coprocessor
|
|
giving memory-mapped streaming the real SNES never had. Their budget: 512 tiles x
|
|
32 bytes (4bpp 8x8) + tilemap ~= 18 KB/frame => **~430 KB/s** at 23.976fps.
|
|
That's *higher* than the 365 KB/s we'd reject on SASI. (ESTIMATE: my arithmetic on
|
|
their stated tile budget, not a measured figure.)
|
|
|
|
Where we're ahead: 256 simultaneous colors from a 65536 palette vs their 4bpp
|
|
sub-palettes needing a tile-aware palette optimizer plus a spatial smoothing pass
|
|
to hide 8x8 palette seams. That problem doesn't exist for us. Plus 68000@10MHz
|
|
vs 65816@3.58MHz, and 2MB vs 128KB.
|
|
|
|
**Most valuable thing in that repo is NOT the codec — it's `data/events/`:**
|
|
516 chapter definitions across 29 scenes as XML, plus
|
|
`data/chapter_event_inventory.md`. That's the arcade scene graph and input-timing
|
|
structure, entirely hardware-independent — the whole game-logic layer we'd
|
|
otherwise reverse-engineer from the arcade ROM.
|
|
|
|
**TODO: check their license before planning to reuse it.**
|
|
Their 516 chapters are finer-grained than our 224 Blu-ray streams, so mapping
|
|
their event table onto our footage means subdividing streams by timecode.
|
|
|
|
Caveat: all of the above is from README/repo-tree summaries, not their source.
|
|
|
|
---
|
|
---
|
|
|
|
# Findings — session 2 (2026-08-23)
|
|
|
|
## 8. CORRECTION to session 1: halving the framerate does NOT halve the bitrate
|
|
|
|
Session 1 measured 365 KB/s for naive delta+RLE at 24 fps and wrote
|
|
"(~183 KB/s at 12fps)". **That extrapolation is wrong.** Decimating to 12 fps
|
|
roughly doubles the per-frame delta, so the *rate* stays nearly flat.
|
|
|
|
Re-measured directly on 12 fps decimated frames (4 scenes, 66 frames):
|
|
|
|
| codec (all LOSSLESS w.r.t. the 256-colour frame) | B/frame | KB/s @12 | 22 min | ratio |
|
|
|---|---|---|---|---|
|
|
| raw 8bpp 256x192 | 49152 | 576 | 743 MB | 1.0:1 |
|
|
| session 1 row-span + RLE | 29055 | 340 | 439 MB | 1.7:1 |
|
|
| XOR vs prev + deflate | 30196 | 354 | 456 MB | 1.6:1 |
|
|
| **changed-spans + deflate** | **21110** | **247** | **319 MB** | **2.3:1** |
|
|
| changed-spans + LZMA | 18759 | 220 | 283 MB | 2.6:1 |
|
|
|
|
Session 1's own RLE re-measured at 12 fps gives **340 KB/s, not 183**.
|
|
Any plan that assumed 183 KB/s was based on a bad number.
|
|
|
|
Deflate-class entropy coding on top of the span payload is worth **1.4x** over
|
|
hand-rolled RLE, and LZ decode is cheap on a 68000 (byte copies), so the
|
|
lossless floor is ~247 KB/s / 319 MB. That is **infeasible on SASI** and
|
|
**tight but real on SCSI**.
|
|
|
|
## 9. Flat 4x4 VQ at k=256 is NOT acceptable — confirmed by eye
|
|
|
|
The risk flagged in 6 is real. At k=256, 4x4:
|
|
|
|
| scene | palette-only PSNR | after VQ | VQ loss |
|
|
|---|---|---|---|
|
|
| 00010 | 38.35 | 29.68 | 8.67 dB |
|
|
| 00020 | 39.90 | 32.67 | 7.22 dB |
|
|
| 00146 | 35.25 | 29.35 | 5.89 dB |
|
|
| 00181 | 41.92 | 32.87 | 9.05 dB |
|
|
|
|
Visually: Dirk's face disintegrates, teeth and eyes turn to mush, ink outlines
|
|
break into 4-pixel stair-steps, colour bleeds across block boundaries.
|
|
|
|

|
|
*Left: 1080p source. Middle: 256-colour palettised 256x192 — the quality ceiling,
|
|
and it is excellent. Right: flat 4x4 VQ at k=256. This is the result that killed
|
|
the flat-VQ architecture.*
|
|
|
|
**Crucially, the 256-colour palettised frame itself looks excellent.** Flat cel
|
|
art with a per-scene median-cut palette and no dithering is near-transparent
|
|
(35-42 dB). So the palette is not the problem and 256 colours is not the
|
|
problem — **block VQ is**. The quality ceiling we should hold ourselves to is
|
|
the palettised frame, not the 1080p source.
|
|
|
|
## 10. Hybrid VQ (Cinepak V1/V4 + SKIP) — this is the codec
|
|
|
|
Per 4x4 block, choose by rate-distortion: SKIP (reuse previous frame),
|
|
V1 (one 4x4 codeword, 1 byte), or V4 (four 2x2 codewords, 4 bytes),
|
|
with a 2-bit-per-block mode header. `lam` is the lagrangian rate knob.
|
|
|
|
Measured, k1=k4=256, 4 scenes (mean of the per-scene table in the session log):
|
|
|
|
| lam | PSNR | loss vs palette | SKIP% | V1% | V4% | B/frame | KB/s @12 |
|
|
|---|---|---|---|---|---|---|---|
|
|
| 0 (max quality) | 33.9 | 4.9 | 30.8 | 18.5 | 50.8 | 7574 | 88.8 |
|
|
| 200 | 31.9 | 5.9 | 44.0 | 37.6 | 18.4 | 4183 | 49.0 |
|
|
| 1000 | 31.6 | 7.3 | 47.4 | 47.7 | 4.9 | 2841 | 33.3 |
|
|
| 5000 | 25.5 | 13.3 | 55.6 | 44.4 | 0.0 | 2134 | 25.0 |
|
|
|
|
At a **matched ~30 KB/s** the hybrid beats flat 4x4 VQ by ~1 dB, and unlike flat
|
|
VQ it keeps scaling: at 89 KB/s it reaches within **4.9 dB of the palette
|
|
ceiling**, which flat VQ cannot reach at any bitrate.
|
|
|
|
Note V4% collapses to 0 at lam=5000 — that is the knob doing exactly what it
|
|
should: under a hard ceiling, detail blocks are the first thing sacrificed.
|
|
|
|
## 11. Codebook size sweep (flat 4x4, for reference)
|
|
|
|
| block | k | PSNR | loss | key B | changed% | KB/s @12 | codebook RAM |
|
|
|---|---|---|---|---|---|---|---|
|
|
| 4x4 | 256 | 30.46 | 8.39 | 3072 | 52.7 | 28.5 | 8K |
|
|
| 4x4 | 1024 | 32.89 | 5.96 | 3840 | 56.6 | 35.6 | 32K |
|
|
|
|
+2.4 dB for 24K more RAM and 7 KB/s. With 2 MB of RAM, a 1024-entry codebook is
|
|
cheap and clearly worth it. (RAM figure is the word-expanded form the blitter
|
|
wants: k * 16 px * 2 bytes.)
|
|
|
|
## 12. Source framing — OPEN
|
|
|
|
The Blu-ray is **full-frame 1920x1080 16:9 with no pillarboxing**. The arcade
|
|
original is 4:3. The extractor currently centre-crops 1440x1080, which is the
|
|
arcade-faithful choice but discards image the 2006 remaster added. Options are
|
|
`crop` (default), `squash`, `wide` in `tools/encoder/extract.py`.
|
|
**Not yet decided; needs an eyeball comparison against arcade reference.**
|
|
|
|
## 13. Stream inventory correction
|
|
|
|
Session 1 said "typical scene clip ~60s". Sampled directly: the ~3-5 MB streams
|
|
are **1.2-1.7 s** clips — these are the individual arcade death/action moments,
|
|
which is exactly the granularity the game logic needs. Some 60 s streams
|
|
(e.g. 00203) are **menu screens, not content**. Any survey must classify
|
|
menu vs content before averaging, or the bitrate numbers are diluted by static
|
|
menus.
|
|
|
|
## 14. A FOURTH false-good result — and the correction
|
|
|
|
Add this to the 4 list. The mechanism was new but the shape was identical.
|
|
|
|
**The false result:** flat and hybrid VQ both showed **+2.4 dB for k=1024 over
|
|
k=256** at an apparently similar bitrate, which made a 1024-entry codebook look
|
|
like an obvious win. The k=1024 quality ladder rendered from that run looked
|
|
great at "45 KB/s".
|
|
|
|
**The bug:** the rate-distortion model in `vq_hybrid.encode()` charged **1 byte**
|
|
per codebook index unconditionally. A 1024-entry codebook needs a **10-bit index,
|
|
stored as 2 bytes**. So every k=1024 measurement understated the V1 and V4
|
|
payload by exactly 2x, *and* the lagrangian mode decision was choosing V4 on the
|
|
belief that four codewords cost 4 bytes when they cost 8.
|
|
|
|
**After charging the true index cost** (`idx_bytes` is now explicit and defaults
|
|
from the codebook size), matched-bitrate comparison on scene 00020:
|
|
|
|
| KB/s | k=256 (1-byte idx) | k=1024 (2-byte idx) |
|
|
|---|---|---|
|
|
| ~32-42 | **33.87 dB** @ 32.5 | 28.91 dB @ 42.3 |
|
|
| ~44-52 | **34.80 dB** @ 44.1 | 35.13 dB @ 52.5 |
|
|
| ~72-86 | **35.87 dB** @ 72.2 | 36.51 dB @ 86.0 |
|
|
|
|
k=1024 buys +0.3 to +0.6 dB for +19% bitrate — a wash at best — and at the low
|
|
end where the SASI profile lives it is **5 dB worse**, because the 2-byte index
|
|
floor dominates once V4 is priced out.
|
|
|
|
**k=256 with 1-byte indices is the shipping choice.** It is also the better
|
|
decoder: a plain `move.b` index with no alignment case, and an 8 KB codebook
|
|
instead of 32 KB.
|
|
|
|
**The general lesson, again:** the comparison was not wrong about VQ, it was
|
|
wrong about *cost*. When a knob looks like a free win, check that the rate model
|
|
is charging for it. Same failure family as 4.1-4.3: a plausible number produced
|
|
by a pipeline that was not measuring what it claimed to measure.
|
|
|
|
## 15. Rate-distortion curve of the shipping codec (k=256, corrected)
|
|
|
|
Scene 00020 (Dirk screaming, close-up face — the hardest case for linework),
|
|
and 00146. Includes the 2-bit-per-block mode header. No entropy coding yet.
|
|
|
|
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | SKIP | V1 | V4 | RAW |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~37% | ~24% | ~13% | ~26% |
|
|
| 100 | 35.87 | 72.2 | 29.04 | 72.5 | ~41% | ~34% | ~21% | ~4% |
|
|
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | ~44% | ~42% | ~14% | 0% |
|
|
| 800 | 33.87 | 32.5 | 27.77 | 36.1 | ~46% | ~48% | ~5% | 0% |
|
|
| 2000 | 27.57 | 25.5 | 24.88 | 30.2 | ~50% | ~49% | ~1% | 0% |
|
|
|
|
Palette ceilings: 00020 = 39.90 dB, 00146 = 35.25 dB.
|
|
|
|

|
|
*The shipping codec across the rate knob. Top: source, palette ceiling, lam=25.
|
|
Bottom: lam=100 (`scsi` profile), lam=300 (`sasi` profile), lam=800.
|
|
Both shipping profiles hold Bluth's linework; the failure only starts past lam=800.*
|
|
|
|
Two things to read off this table:
|
|
- **The cliff is between lam=800 and lam=2000.** That is where V4 is priced out
|
|
entirely and detail blocks have nowhere to go. Do not ship past lam~800.
|
|
- **RAW is doing real work at high bitrate** (26% of blocks at lam=25) and
|
|
vanishes by lam=300. It is what makes the top of the curve reach the palette
|
|
ceiling, and it costs the decoder nothing — RAW is the cheapest mode to blit.
|
|
|
|
## 16. Licences cleared for the game-logic layer
|
|
|
|
Both checked this session:
|
|
|
|
- **astrobleem/SNES-SuperDragonsLairArcade — MIT**, "Copyright (c) 2026 Chad
|
|
Doebelin". `data/events/` holds 516 XML chapter definitions with timing and
|
|
event data. Reusable with attribution.
|
|
- **icculus/DirkSimple — zlib.** Independent from-scratch reimplementation of
|
|
the game logic in Lua, scene/timing tables in `game.lua`. Also permissive.
|
|
|
|
Having **two independent permissively-licensed transcriptions** of the arcade
|
|
scene graph is better than one: they can be diffed against each other to catch
|
|
transcription errors before any of it is committed to 68000 tables.
|
|
|
|
## 17. The profiles were set far too low — and entropy coding is a CPU trap
|
|
|
|
Prompted by the user asking why the SCSI profile was only 75 KB/s. It should not
|
|
have been. Two separate errors, one of them serious.
|
|
|
|
### 17.1 The profile bitrates were not derived from the hardware at all
|
|
|
|
They were read off the knee of the rate-distortion curve and then presented as
|
|
though bandwidth-driven. Against the (folklore) bus figures from 5:
|
|
|
|
| profile | was | bus figure | utilisation |
|
|
|---|---|---|---|
|
|
| `sasi` | 45 KB/s | ~300-500 KB/s | **12%** |
|
|
| `scsi` | 75 KB/s | ~1 MB/s | **7%** |
|
|
|
|
Nothing justified leaving 90% of the pipe unused. Raised to `sasi` 110 KB/s
|
|
(lam=60) and `scsi` 280 KB/s (lam=10), which is 35% and 28% utilisation —
|
|
still conservative, because the bus figures are folklore.
|
|
|
|
### 17.2 CPU is NOT the reason to stay low — but entropy coding would be
|
|
|
|
Budget is 833,333 cycles/frame at 12 fps. At session 1's measured ~6.5 cycles
|
|
per GVRAM pixel write:
|
|
|
|
| work | cycles | % of budget |
|
|
|---|---|---|
|
|
| blit 20.1% of pixels (session 1's 24fps figure) | 64k | 7.7% |
|
|
| blit 40% of pixels (the same content at 12fps) | 128k | 15.3% |
|
|
| **blit the FULL frame, every frame** | **319k** | **38.3%** |
|
|
| deflate decode, ~30 KB/frame output | 1,800k | **216%** |
|
|
| LZ4/LZSS decode, ~30 KB/frame output | 450k | **54%** |
|
|
|
|
Two conclusions, and the second one corrects 8:
|
|
|
|
- **Raising the VQ bitrate is nearly free on CPU.** Even a full-frame pixel-exact
|
|
blit fits in 38% of budget, and VQ decode is table copies — RAW, the mode that
|
|
dominates at high bitrate, is the *cheapest* mode to blit, not the dearest.
|
|
- **The 247 KB/s "lossless changed-spans + deflate" figure in 8 is a compression
|
|
upper bound, NOT a shippable design.** Deflate's Huffman decode is bitwise and
|
|
costs about 2.2x the entire frame budget on a 68000. Even byte-oriented LZ4 at
|
|
54% leaves too little beside a 38% blit. **Do not plan on entropy coding.**
|
|
All profile bitrates are raw payload.
|
|
|
|
This inverts session 1's "the CPU is idle, I/O is the ceiling" for the *decode*
|
|
path specifically: the blit is cheap, but any bit-oriented decompressor is not.
|
|
VQ is the right architecture precisely because its decode is a table copy.
|
|
|
|
### 17.3 The hybrid at lam=0 IS the lossless codec
|
|
|
|
Measured, un-entropy-coded raw payload, and deflated for reference only:
|
|
|
|
| scene | lam=0 raw | lam=0 deflated | lossless changed-spans+deflate | PSNR |
|
|
|---|---|---|---|---|
|
|
| 00020 | 442.1 KB/s | 274.5 KB/s | 267.3 KB/s | 39.90 = ceiling |
|
|
| 00146 | 467.6 KB/s | 223.2 KB/s | 219.1 KB/s | 35.25 = ceiling |
|
|
|
|
The hybrid at `lam=0` converges to within 3% of the purpose-built lossless coder.
|
|
That confirms the architecture unifies: there is no separate lossless path to
|
|
maintain, just the same bitstream with the knob open.
|
|
|
|
### 17.4 Full curve in raw (shippable) bytes
|
|
|
|
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | RAW% |
|
|
|---|---|---|---|---|---|
|
|
| 0 | **39.90** (exact) | 442.1 | **35.25** (exact) | 467.6 | ~76% |
|
|
| 10 | 39.38 | 248.1 | 32.27 | 305.2 | ~44% |
|
|
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~26% |
|
|
| 60 | 36.94 | 108.0 | 29.61 | 103.1 | ~10% |
|
|
| 150 | 35.31 | 55.6 | 28.63 | 56.1 | ~1% |
|
|
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | 0% |
|
|
|
|
### 17.5 This makes the blocked disk benchmark critical-path
|
|
|
|
Session 1 judged it "NOT on the critical path" because VQ at 30 KB/s was correct
|
|
whether SASI did 300 or 600 KB/s. That reasoning no longer holds. The profiles
|
|
now sit at 110 and 280 KB/s, close enough to the folklore ceilings that the
|
|
error bars matter, and **if SCSI sustains >=800 KB/s the correct `scsi` profile
|
|
is lam=0 — pixel-exact video.** Whether this port ships transparent or lossy on
|
|
SCSI is now waiting on one measurement.
|
|
|
|
## 18. Peak-to-mean burstiness — the mean was hiding the problem
|
|
|
|
Prompted by the user clarifying that the bandwidth figure is **4 Mbps = 488 KB/s**,
|
|
not 4 MB/s. That is ~8x tighter than what 17 was reasoning against, and it
|
|
changes the answer.
|
|
|
|
Per-frame instantaneous rate (video + 7.8 KB/s audio), 12 fps:
|
|
|
|
| scene | lam | mean | p90 | **max** | peak/mean | max as % of 488 KB/s |
|
|
|---|---|---|---|---|---|---|
|
|
| 00010 | 60 | 95.0 | 127.3 | 138.8 | 1.46 | 28.4% |
|
|
| 00010 | 10 | 198.9 | 266.1 | 284.0 | 1.43 | 58.2% |
|
|
| 00020 | 60 | 115.8 | 155.4 | 222.3 | 1.92 | 45.5% |
|
|
| 00020 | 10 | 255.9 | 391.2 | **470.8** | 1.84 | **96.4%** |
|
|
|
|
**The `scsi` profile as committed in f0f2f80 does not fit 4 Mbps.** Its mean is a
|
|
comfortable 52% of the pipe, but it peaks at 96.4% — and a frame that arrives
|
|
late is a *dropped frame*, not a slow one. Sizing a real-time stream on the mean
|
|
is the mistake; peak/mean is 1.4-1.9x on 1.2-1.7s clips and will be worse across
|
|
a full scene.
|
|
|
|
Two ways out, and only one is good:
|
|
- Size for the peak: `lam=25`, mean 194 KB/s. Costs a full step of quality.
|
|
- **Rate-control to the mean and carry a leaky bucket:** `lam=10` fits, and buys
|
|
back +0.7 dB (00020) / +1.2 dB (00146).
|
|
|
|
`ratectl.py` was written in session 2 but **never wired into `encode.py`**. This
|
|
demotes that from a loose end to the highest-value unfinished work in the repo.
|
|
|
|
## 19. Cycle-stealing DMA is not free DMA — 5 was wrong
|
|
|
|
FINDINGS 5 concluded "because it's DMA, streaming costs essentially no CPU —
|
|
this stacks with the 8% blit utilisation. The 68000 really is nearly idle."
|
|
|
|
The HD63450 steals bus cycles from the 68000 at roughly 8 clocks per 16-bit word:
|
|
|
|
| stream | words/s | clocks/s | CPU stolen | + full-frame blit |
|
|
|---|---|---|---|---|
|
|
| 110 KB/s | 56,320 | 450,560 | 4.5% | 42.8% |
|
|
| 250 KB/s | 128,000 | 1,024,000 | 10.2% | 48.5% |
|
|
| 450 KB/s | 230,400 | 1,843,200 | 18.4% | 56.7% |
|
|
| 488 KB/s | 249,856 | 1,998,848 | 20.0% | 58.3% |
|
|
|
|
At the rates the profiles now use, streaming costs **10-20% of the machine**.
|
|
Still affordable — nothing here breaks — but **bandwidth and CPU are one budget,
|
|
not two**, and any future headroom argument has to spend from both. The
|
|
"nearly idle" framing should not be reused.
|
|
|
|
(The 8 clocks/word figure is session 1's ESTIMATE from HD63450 timing, not a
|
|
measurement. It is the weakest link in this table.)
|
|
|
|
## 20. Where the profiles should come from
|
|
|
|
`tools/encoder/profile_gen.py` now derives lam from a bandwidth figure rather
|
|
than from the shape of the RD curve, accounting for audio, peak/mean, and
|
|
reporting DMA steal. Full benchmarking methodology — and why MAME cannot answer
|
|
the bandwidth question — is in `docs/BENCHMARK.md`.
|
|
|
|
The 4 Mbps figure itself is **user-supplied and its provenance is not recorded**.
|
|
Every profile now hangs off it, so it is worth pinning down.
|