Session 2: hybrid VQ codec, two quality profiles, three corrections

Answers session 1's critical-path question. Flat 4x4 VQ at k=256 was prototyped
and REJECTED by eye: Dirk's face disintegrates and ink outlines break into
4-pixel stair-steps. The 256-colour palettised frame is excellent, so the
palette was never the problem -- block VQ was.

Replaced it with a Cinepak-style hybrid: each 4x4 block is SKIP, one 4x4
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
rate-distortion. The RAW escape makes lam=0 pixel-exact (measured 0.00 dB loss),
so the quality knob spans lossless to heavily-compressed in one bitstream.

Per the user's decision, ships TWO quality profiles from that one codec, one
decoder and one bitstream -- only the rate knob differs:
  sasi  45 KB/s  lam=300  34.8 dB   stock 10MHz ACE/EXPERT
  scsi  75 KB/s  lam=100  35.9 dB   Super/XVI or CZ-6BS1

Three corrections to earlier numbers:

1. Session 1's "183 KB/s at 12fps" was a bad extrapolation. Halving the
   framerate does not halve the bitrate -- decimation roughly doubles the
   per-frame delta. Re-measured directly: 340 KB/s for session 1's own RLE,
   247 KB/s for changed-spans+deflate. The lossless floor is 319 MB.

2. A FOURTH false-good result, same family as the three in FINDINGS 4:
   k=1024 codebooks appeared to buy +2.4 dB free, because the rate model
   charged 1 byte for a 10-bit index. Charging the true cost reverses the
   verdict -- k=256 wins at every matched bitrate, and by 5 dB at the low end
   where the SASI profile lives. k=256 ships.

3. Stream inventory: the ~3-5MB clips are 1.2-1.7s, not ~60s, and some 60s
   streams are menus, not content. Any survey must classify before averaging.

Also cleared both candidate sources for the game-logic layer: the SNES project
is MIT and DirkSimple is zlib, so the arcade scene graph can be imported and
the two transcriptions diffed against each other.

Encoder is working end-to-end: extract.py -> vq/vq_hybrid/ratectl -> encode.py,
emitting a big-endian DLX1 container the 68000 can parse with plain moves.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
prosolis
2026-08-23 11:56:08 -07:00
parent 65112b9305
commit e4062ed294
11 changed files with 917 additions and 104 deletions
+3
View File
@@ -8,3 +8,6 @@ assets/audio/
assets/frames/ assets/frames/
build/ build/
roms/ roms/
__pycache__/
*.pyc
*.dlx
+16 -1
View File
@@ -19,9 +19,24 @@ docs/ findings, status, hardware reference
tools/analysis/ frame-analysis scripts (01/02 marked BROKEN as regression refs) tools/analysis/ frame-analysis scripts (01/02 marked BROKEN as regression refs)
tools/bench/ MAME Lua injection harness + 68000 benchmark sources tools/bench/ MAME Lua injection harness + 68000 benchmark sources
tools/vasm/ vasm m68k assembler (built from source) tools/vasm/ vasm m68k assembler (built from source)
tools/encoder/ VQ encoder (not yet written) tools/encoder/ hybrid VQ encoder + DLX1 container writer (working)
src/player/ 68000 player (not yet written) src/player/ 68000 player (not yet written)
assets/ extracted frames/audio (gitignored) assets/ extracted frames/audio (gitignored)
``` ```
## Encoder
```
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile sasi --preview p.png
```
Two quality profiles ship from one codec and one decoder — `sasi` (45 KB/s) and
`scsi` (120 KB/s) are two points on the same rate-distortion curve. The codec is
a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four
2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion.
The RAW escape means `lam=0` is pixel-exact against the palettised frame, so the
quality knob spans lossless to heavily-compressed without changing the bitstream.
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored — supply your own. Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored — supply your own.
+182
View File
@@ -188,3 +188,185 @@ Their 516 chapters are finer-grained than our 224 Blu-ray streams, so mapping
their event table onto our footage means subdividing streams by timecode. their event table onto our footage means subdividing streams by timecode.
Caveat: all of the above is from README/repo-tree summaries, not their source. Caveat: all of the above is from README/repo-tree summaries, not their source.
---
---
# Findings — session 2 (2026-08-23)
## 8. CORRECTION to session 1: halving the framerate does NOT halve the bitrate
Session 1 measured 365 KB/s for naive delta+RLE at 24 fps and wrote
"(~183 KB/s at 12fps)". **That extrapolation is wrong.** Decimating to 12 fps
roughly doubles the per-frame delta, so the *rate* stays nearly flat.
Re-measured directly on 12 fps decimated frames (4 scenes, 66 frames):
| codec (all LOSSLESS w.r.t. the 256-colour frame) | B/frame | KB/s @12 | 22 min | ratio |
|---|---|---|---|---|
| raw 8bpp 256x192 | 49152 | 576 | 743 MB | 1.0:1 |
| session 1 row-span + RLE | 29055 | 340 | 439 MB | 1.7:1 |
| XOR vs prev + deflate | 30196 | 354 | 456 MB | 1.6:1 |
| **changed-spans + deflate** | **21110** | **247** | **319 MB** | **2.3:1** |
| changed-spans + LZMA | 18759 | 220 | 283 MB | 2.6:1 |
Session 1's own RLE re-measured at 12 fps gives **340 KB/s, not 183**.
Any plan that assumed 183 KB/s was based on a bad number.
Deflate-class entropy coding on top of the span payload is worth **1.4x** over
hand-rolled RLE, and LZ decode is cheap on a 68000 (byte copies), so the
lossless floor is ~247 KB/s / 319 MB. That is **infeasible on SASI** and
**tight but real on SCSI**.
## 9. Flat 4x4 VQ at k=256 is NOT acceptable — confirmed by eye
The risk flagged in 6 is real. At k=256, 4x4:
| scene | palette-only PSNR | after VQ | VQ loss |
|---|---|---|---|
| 00010 | 38.35 | 29.68 | 8.67 dB |
| 00020 | 39.90 | 32.67 | 7.22 dB |
| 00146 | 35.25 | 29.35 | 5.89 dB |
| 00181 | 41.92 | 32.87 | 9.05 dB |
Visually: Dirk's face disintegrates, teeth and eyes turn to mush, ink outlines
break into 4-pixel stair-steps, colour bleeds across block boundaries.
![flat 4x4 VQ failure](images/flat_vq_failure_00010.png)
*Left: 1080p source. Middle: 256-colour palettised 256x192 — the quality ceiling,
and it is excellent. Right: flat 4x4 VQ at k=256. This is the result that killed
the flat-VQ architecture.*
**Crucially, the 256-colour palettised frame itself looks excellent.** Flat cel
art with a per-scene median-cut palette and no dithering is near-transparent
(35-42 dB). So the palette is not the problem and 256 colours is not the
problem — **block VQ is**. The quality ceiling we should hold ourselves to is
the palettised frame, not the 1080p source.
## 10. Hybrid VQ (Cinepak V1/V4 + SKIP) — this is the codec
Per 4x4 block, choose by rate-distortion: SKIP (reuse previous frame),
V1 (one 4x4 codeword, 1 byte), or V4 (four 2x2 codewords, 4 bytes),
with a 2-bit-per-block mode header. `lam` is the lagrangian rate knob.
Measured, k1=k4=256, 4 scenes (mean of the per-scene table in the session log):
| lam | PSNR | loss vs palette | SKIP% | V1% | V4% | B/frame | KB/s @12 |
|---|---|---|---|---|---|---|---|
| 0 (max quality) | 33.9 | 4.9 | 30.8 | 18.5 | 50.8 | 7574 | 88.8 |
| 200 | 31.9 | 5.9 | 44.0 | 37.6 | 18.4 | 4183 | 49.0 |
| 1000 | 31.6 | 7.3 | 47.4 | 47.7 | 4.9 | 2841 | 33.3 |
| 5000 | 25.5 | 13.3 | 55.6 | 44.4 | 0.0 | 2134 | 25.0 |
At a **matched ~30 KB/s** the hybrid beats flat 4x4 VQ by ~1 dB, and unlike flat
VQ it keeps scaling: at 89 KB/s it reaches within **4.9 dB of the palette
ceiling**, which flat VQ cannot reach at any bitrate.
Note V4% collapses to 0 at lam=5000 — that is the knob doing exactly what it
should: under a hard ceiling, detail blocks are the first thing sacrificed.
## 11. Codebook size sweep (flat 4x4, for reference)
| block | k | PSNR | loss | key B | changed% | KB/s @12 | codebook RAM |
|---|---|---|---|---|---|---|---|
| 4x4 | 256 | 30.46 | 8.39 | 3072 | 52.7 | 28.5 | 8K |
| 4x4 | 1024 | 32.89 | 5.96 | 3840 | 56.6 | 35.6 | 32K |
+2.4 dB for 24K more RAM and 7 KB/s. With 2 MB of RAM, a 1024-entry codebook is
cheap and clearly worth it. (RAM figure is the word-expanded form the blitter
wants: k * 16 px * 2 bytes.)
## 12. Source framing — OPEN
The Blu-ray is **full-frame 1920x1080 16:9 with no pillarboxing**. The arcade
original is 4:3. The extractor currently centre-crops 1440x1080, which is the
arcade-faithful choice but discards image the 2006 remaster added. Options are
`crop` (default), `squash`, `wide` in `tools/encoder/extract.py`.
**Not yet decided; needs an eyeball comparison against arcade reference.**
## 13. Stream inventory correction
Session 1 said "typical scene clip ~60s". Sampled directly: the ~3-5 MB streams
are **1.2-1.7 s** clips — these are the individual arcade death/action moments,
which is exactly the granularity the game logic needs. Some 60 s streams
(e.g. 00203) are **menu screens, not content**. Any survey must classify
menu vs content before averaging, or the bitrate numbers are diluted by static
menus.
## 14. A FOURTH false-good result — and the correction
Add this to the 4 list. The mechanism was new but the shape was identical.
**The false result:** flat and hybrid VQ both showed **+2.4 dB for k=1024 over
k=256** at an apparently similar bitrate, which made a 1024-entry codebook look
like an obvious win. The k=1024 quality ladder rendered from that run looked
great at "45 KB/s".
**The bug:** the rate-distortion model in `vq_hybrid.encode()` charged **1 byte**
per codebook index unconditionally. A 1024-entry codebook needs a **10-bit index,
stored as 2 bytes**. So every k=1024 measurement understated the V1 and V4
payload by exactly 2x, *and* the lagrangian mode decision was choosing V4 on the
belief that four codewords cost 4 bytes when they cost 8.
**After charging the true index cost** (`idx_bytes` is now explicit and defaults
from the codebook size), matched-bitrate comparison on scene 00020:
| KB/s | k=256 (1-byte idx) | k=1024 (2-byte idx) |
|---|---|---|
| ~32-42 | **33.87 dB** @ 32.5 | 28.91 dB @ 42.3 |
| ~44-52 | **34.80 dB** @ 44.1 | 35.13 dB @ 52.5 |
| ~72-86 | **35.87 dB** @ 72.2 | 36.51 dB @ 86.0 |
k=1024 buys +0.3 to +0.6 dB for +19% bitrate — a wash at best — and at the low
end where the SASI profile lives it is **5 dB worse**, because the 2-byte index
floor dominates once V4 is priced out.
**k=256 with 1-byte indices is the shipping choice.** It is also the better
decoder: a plain `move.b` index with no alignment case, and an 8 KB codebook
instead of 32 KB.
**The general lesson, again:** the comparison was not wrong about VQ, it was
wrong about *cost*. When a knob looks like a free win, check that the rate model
is charging for it. Same failure family as 4.1-4.3: a plausible number produced
by a pipeline that was not measuring what it claimed to measure.
## 15. Rate-distortion curve of the shipping codec (k=256, corrected)
Scene 00020 (Dirk screaming, close-up face — the hardest case for linework),
and 00146. Includes the 2-bit-per-block mode header. No entropy coding yet.
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | SKIP | V1 | V4 | RAW |
|---|---|---|---|---|---|---|---|---|
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~37% | ~24% | ~13% | ~26% |
| 100 | 35.87 | 72.2 | 29.04 | 72.5 | ~41% | ~34% | ~21% | ~4% |
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | ~44% | ~42% | ~14% | 0% |
| 800 | 33.87 | 32.5 | 27.77 | 36.1 | ~46% | ~48% | ~5% | 0% |
| 2000 | 27.57 | 25.5 | 24.88 | 30.2 | ~50% | ~49% | ~1% | 0% |
Palette ceilings: 00020 = 39.90 dB, 00146 = 35.25 dB.
![quality ladder](images/quality_ladder_00020.png)
*The shipping codec across the rate knob. Top: source, palette ceiling, lam=25.
Bottom: lam=100 (`scsi` profile), lam=300 (`sasi` profile), lam=800.
Both shipping profiles hold Bluth's linework; the failure only starts past lam=800.*
Two things to read off this table:
- **The cliff is between lam=800 and lam=2000.** That is where V4 is priced out
entirely and detail blocks have nowhere to go. Do not ship past lam~800.
- **RAW is doing real work at high bitrate** (26% of blocks at lam=25) and
vanishes by lam=300. It is what makes the top of the curve reach the palette
ceiling, and it costs the decoder nothing — RAW is the cheapest mode to blit.
## 16. Licences cleared for the game-logic layer
Both checked this session:
- **astrobleem/SNES-SuperDragonsLairArcade — MIT**, "Copyright (c) 2026 Chad
Doebelin". `data/events/` holds 516 XML chapter definitions with timing and
event data. Reusable with attribution.
- **icculus/DirkSimple — zlib.** Independent from-scratch reimplementation of
the game logic in Lua, scene/timing tables in `game.lua`. Also permissive.
Having **two independent permissively-licensed transcriptions** of the arcade
scene graph is better than one: they can be diffed against each other to catch
transcription errors before any of it is committed to 68000 tables.
+104 -103
View File
@@ -1,4 +1,4 @@
# Status & next-session handoff — end of session 1 (2026-08-23) # Status & next-session handoff — end of session 2 (2026-08-23)
## Decisions locked ## Decisions locked
@@ -7,133 +7,134 @@
| Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint | | Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint |
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c | | Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
| Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway | | Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway |
| Codec | 4x4 vector quantization, per-scene codebook + block delta | CPU is idle, I/O is the ceiling — spend cycles to buy bandwidth | | **Codec** | **hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion** | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 |
| **Quality modes** | **two: `sasi` and `scsi`** (USER DECISION, session 2) | one codec, one decoder, one bitstream; only `lam` differs |
| Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win | | Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win |
| Medium | SCSI HDD image (.hds) | but see SASI/SCSI split below |
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something | | Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
| SNES project reuse | **MIT — cleared** | `data/events/` scene graph is reusable with attribution |
**OPEN QUESTION for the user:** stock 10MHz machines are **SASI**, not SCSI. ### The SASI/SCSI question is RESOLVED
Three options, not yet chosen: Session 1 left "which machine do we target" open. The user's answer: **ship both**,
1. Stock 10MHz + SASI (purist) — VQ becomes mandatory as two quality profiles. This is now implemented rather than hypothetical — the
2. Stock 10MHz + CZ-6BS1 SCSI board — relieves I/O, keeps CPU honest bitrate ceiling is a build parameter in `tools/encoder/ratectl.py`:
3. Super/XVI baseline — built-in SCSI, still a 10MHz 68000
Recommendation: make the codec's bitrate ceiling a **build parameter**, so one | profile | target | lam | quality (00020 / 00146) | machine |
encoder serves all three and the target is chosen at package time. |---|---|---|---|---|
| `sasi` | 45 KB/s | 300 | 34.8 / 28.3 dB | stock 10MHz ACE/EXPERT |
| `scsi` | 75 KB/s | 100 | 35.9 / 29.0 dB | Super/XVI, or CZ-6BS1 board |
Codebooks are **k=256 with 1-byte indices** in both profiles. k=1024 was measured
and rejected — see FINDINGS 14, it was a false-good result from a rate model
that undercharged the index. Do not ship past `lam~800`; FINDINGS 15 has the cliff.
Because of the RAW escape mode, `lam=0` is **pixel-exact** against the palettised
frame (measured 0.00 dB loss). The profiles are two points on one continuous
rate-distortion curve, not two codecs.
--- ---
## Working setup ## What session 2 settled
**MAME ROMs**`~/mame/roms/x68000.zip` (present, working). 1. **The critical-path question is answered.** "Does VQ soften Bluth's linework
Must pass **`-bios ipl10`**; the default BIOS is `cz600ce`, whose split unacceptably?" — **flat 4x4 k=256 VQ: yes, badly. Hybrid VQ with k=1024: no.**
even/odd IPL halves (`rh-ix0897cezz.ic12` / `rh-ix0898cezz.ic11`) are absent. Verified by eye, not just PSNR. See `docs/FINDINGS.md` 9-11.
`-verifyroms` will still report those two as missing — this is expected and harmless. 2. **Session 1's 12fps bitrate was wrong** (183 KB/s claimed, 340 KB/s measured).
Halving the framerate does not halve the bitrate. FINDINGS 8.
2b. **A fourth false-good result was produced and caught this session** — k=1024
codebooks looked like a +2.4 dB free win because the rate model charged 1 byte
for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
3. **The 256-colour palettised frame is the real quality ceiling** and it looks
excellent. Judge the codec against that, not against 1080p.
4. Encoder exists and produces a real bitstream: `tools/encoder/`.
Boots headless at ~430-480% speed: ---
## Encoder — working
```
python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
python3 tools/encoder/encode.py /tmp/fr_00020 out.dlx --profile sasi --preview p.png
```
| file | role |
|---|---|
| `extract.py` | .m2ts -> 256x192 PNGs, 12fps, spatial-only denoise |
| `vq.py` | palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR |
| `vq_hybrid.py` | the codec: 4 block modes + lagrangian mode decision |
| `ratectl.py` | SASI/SCSI profiles, leaky-bucket rate control |
| `encode.py` | CLI + `DLX1` container writer |
`DLX1` container layout is documented in the `encode.py` docstring. All
multi-byte fields are **big-endian** so the 68000 reads them with a plain `move`.
### Known encoder gaps
- **Rate control is written but not yet wired into `encode.py`** — the CLI uses a
fixed `lam` from the profile. `ratectl.encode_rate_controlled()` exists and
builds a lam-ladder per frame; it needs hooking up and validating.
- **Payload is not entropy-coded.** Deflate on the payload should buy ~1.4x
(measured on the lossless path, FINDINGS 8). LZ decode is cheap on a 68000.
- Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
- `_paint` is a Python per-block loop — fine for prototyping, slow for a full
disc encode. Vectorise before the 224-stream run.
---
## Working setup (unchanged from session 1, re-verified)
**MAME ROMs**`~/mame/roms/x68000.zip`. Must pass **`-bios ipl10`**.
``` ```
mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3 mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3
``` ```
**Assembler**`tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s`
**Assembler** — vasm built from source, binary at `tools/vasm/vasmm68k_mot` **Blu-ray**`udisksctl loop-setup -r -f DRAGONS_LAIR.iso` -> `/media/reala-misaki/BDROM`
(source tarball alongside it). Verified correct 68000 output. (still mounted as of end of session 2).
```
tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s
```
**Blu-ray** — mount with: **MAME Lua harness**`tools/bench/*.lua`, working. Three gotchas (retain the
``` notifier subscription in a global; the stack register is `SP` not `A7`;
udisksctl loop-setup -r -f DRAGONS_LAIR.iso # -> /media/reala-misaki/BDROM `autoboot_script` fires at PC=0 before boot) are documented in FINDINGS.
```
NOTE: this loop mount is still active from session 1. Re-mount if the machine rebooted. **Two shell traps, both hit again this session:**
- piping MAME (or any long job) through `grep` block-buffers — write to a file.
- `pkill -f <pattern>` matches your own shell and kills it (exit 144).
Use `pkill -x` or kill by PID.
--- ---
## MAME Lua harness — WORKING, reusable ## STILL BLOCKED: disk throughput benchmark
`tools/bench/*.lua` inject 68000 machine code straight into emulated RAM and time Unchanged from session 1 — `IOCS _B_READ` returns -1 uniformly. Full diagnosis
it against the emulated clock. No bootable disk or OS required. This is the and the four untested hypotheses are in session 1's notes (git history of this
measurement rig for all future cycle-cost work (blit timing, decoder benchmarks). file, commit 65112b9).
Pattern: **This now matters more than session 1 thought.** Session 1 dismissed it because
``` "VQ at 30 KB/s is correct whether SASI does 300 or 600 KB/s". But we now ship
mame x68000 -bios ipl10 -video none -sound none -nothrottle \ *two profiles*, and the profile bitrates (45 / 120 KB/s) are set against
-seconds_to_run 30 -plugins -autoboot_script yourscript.lua **folklore** bandwidth figures. A real measurement would let us set them
``` honestly instead of conservatively. Next move is the untried SCSI path:
`-exp1 cz6bs1 -hard disk.chd`.
### Three MAME Lua gotchas — all cost real time, all now solved
1. **Retain the notifier subscription.** `emu.add_machine_frame_notifier()` returns
a token; if you drop it into a chunk-local it is garbage-collected and the
callback **silently stops firing**. Assign it to a **global** (`SUB = ...`).
2. **The stack pointer is `SP`, not `A7`** in `cpu.state[...]`.
Full list: A0-A6, D0-D7, PC, SP, SR, USP, CURPC, CURFLAGS, IR.
3. **`autoboot_script` fires at time=0, before boot** (PC=0). Wait until
`machine.time` >= ~5s before injecting, or IOCS is not yet initialised.
Also: piping MAME through `grep` block-buffers output — write raw to a file when
backgrounding, or you will see an empty log and assume a hang.
And never `pkill -f 'mame x68000'` — the pattern matches your own shell and kills it
(exit 144). Use `pkill -x mame`.
---
## BLOCKED: disk throughput benchmark
**Goal:** measure real SASI/SCSI KB/s to replace the folklore figures in FINDINGS.md §5.
**Status:** harness fully working; the IOCS call itself fails.
`IOCS _B_READ` ($46 via `TRAP #15`; d1.hb=PDA, d2.l=position, d3.l=bytes, a1=buffer)
returns **`FFFFFFFF` (-1), zero reads**, uniformly across:
- all 16 PDA values $80-$8F
- both d1 encodings (PDA in bits 31-24 and bits 15-8)
- image sizes 10MB / 20MB / 40MB
The uniformity is the diagnostic: calls are **dispatched and cleanly rejected**,
so `TRAP #15` and IOCS are reachable. MAME does mount the image
(`:x68k_hdc: opened image file bench.hdf`).
**Untested hypotheses, in rough order of likelihood:**
1. The raw image has no X68000 SASI format, so the IPL's boot scan never registered
a usable drive and IOCS refuses. Would need Human68k to format one — **we have
no Human68k image on this system.**
2. MAME's `x68k_hdc` SASI implementation may be too partial for IOCS-level reads.
3. `SP=$8000` may put the injected stack on top of the IOCS work area in low RAM.
Try a much higher stack.
4. The **SCSI path was never tried** — this is the obvious next move and is more
relevant to the target anyway:
`-exp1 cz6bs1 -hard disk.chd` with `exp1:cz6bs1:scsi:0 harddisk`
(`-listmedia` gains a `harddisk` slot accepting .chd/.hd/.hdv/.2mg/.hdi).
**Honest assessment: this benchmark is NOT on the critical path.** The VQ codec
(~30 KB/s) is correct whether SASI does 300 or 600 KB/s. Do not let it block the
encoder. Its real value is deciding whether the *simpler* row-span codec could
have sufficed.
Caveat if resumed: MAME idealizes drive seek latency. That's acceptable because the
realistic deployment is BlueSCSI/SCSI2SD (SD-backed, no mechanical seek), so what
gets measured is the bus/DMAC/controller path — the genuine ceiling. The caveat
only bites for a real period spinning drive.
--- ---
## Next steps, in priority order ## Next steps, in priority order
1. **Build the VQ encoder** (`tools/encoder/`) — 4x4 blocks, per-scene codebook, 1. **Wire rate control into `encode.py`** and validate that the hard ceiling
block delta. Emit sample PNGs for visual evaluation. **The open question is actually holds on an action scene (the whole point of choosing VQ).
whether VQ softens Bluth's ink linework unacceptably — decide by eye before 2. **Entropy-code the payload** (deflate) — ~1.4x for cheap 68000 decode cost.
committing to the architecture.** 3. **68000 decoder skeleton**: parse `DLX1`, expand codebooks to word-per-pixel,
2. **Full-disc survey** — all 224 streams, not 5s samples, to firm up bitrate blit V1/V4/RAW/SKIP. Measure real cycles with the existing MAME Lua harness —
(current numbers are +/-30%) and map streams onto the arcade scene graph. this is the first time the harness gets used for its actual purpose.
3. **Check the SNES project's license**, then evaluate reusing `data/events/` 4. **Full-disc survey** — classify menu vs content first (FINDINGS 13), then
(516 chapters / 29 scenes) as the scene-graph and input-timing layer. measure bitrate across all 224 streams per profile.
4. Resolve the SASI/SCSI target question with the user. 5. **Resolve the framing question** (FINDINGS 12: crop vs squash vs wide).
5. Optionally unblock the disk benchmark via the SCSI path. 6. Unblock the disk benchmark via the SCSI path, then re-set profile bitrates.
6. 68000 player skeleton: CRTC init for 256x192x256c, `movem.l` blitter, 7. Import the SNES project's `data/events/` (MIT, cleared) as the scene graph.
ADPCM via HD63450 DMA. Cross-check against DirkSimple (zlib) which has the same data independently.
8. ADPCM audio: MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in `ratectl`,
not yet extracted or encoded.
## Not yet started ## Not yet started
- Any 68000 player code - Any 68000 player code
- ADPCM audio extraction/encoding (MSM6258, 15.6kHz mono, ~7.8 KB/s, ~10MB for 22min) - ADPCM audio extraction/encoding
- Disk image packaging / container format - Disk image packaging
- Game logic (scene branching, input windows, death clips) - Game logic (scene branching, input windows, death clips)
Binary file not shown.

After

Width:  |  Height:  |  Size: 680 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.3 MiB

+148
View File
@@ -0,0 +1,148 @@
#!/usr/bin/env python3
"""Encode one scene to the DLX bitstream, at a chosen quality profile.
python3 tools/encoder/encode.py <frames_dir> <out.dlx> [--profile sasi|scsi]
[--lam N] [--fps 12] [--preview out.png]
Container (little-endian is WRONG here -- the 68000 is big-endian, so every
multi-byte field is big-endian and the decoder can read it with a plain move.w):
header, 32 bytes
0 'DLX1' magic
4 u16 width, u16 height
8 u16 fps, u16 nframes
12 u16 k1, u16 k4 codebook sizes
16 u32 palette offset (256 * 3 bytes, RGB888 -- the player converts
to the X68000's GRB555 at load time)
20 u32 cb1 offset (k1 * 16 bytes of palette indices)
24 u32 cb4 offset (k4 * 4 bytes)
28 u32 frames offset
then, per frame:
u32 payload length, then
ceil(nblocks*2/8) bytes of 2-bit mode headers, MSB-first, block raster order
then payloads in block order: V1 -> 1 byte, V4 -> 4 bytes, RAW -> 16 bytes
Codebooks are emitted as palette INDICES, not pixels. The player expands them
once at load time into word-per-pixel form so the blitter can movem them
straight into GVRAM -- k1=1024 costs 1024*16*2 = 32 KB of the 2 MB.
"""
import argparse, struct, sys, os
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import numpy as np
import vq as VQ, vq_hybrid as H, ratectl as RC
def pack_modes(mode):
"""2 bits per block, MSB-first -- cheap for the 68000 to shift out."""
n = len(mode)
out = bytearray((n * 2 + 7) // 8)
for i, m in enumerate(mode):
out[i // 4] |= (int(m) & 3) << (6 - 2 * (i % 4))
return bytes(out)
def frame_payload(mode, l1, l4g, src_idx, nbx):
body = bytearray()
for b, mo in enumerate(mode):
if mo == 1:
body += _idx(l1[b])
elif mo == 2:
for j in range(4):
body += _idx(l4g[b][j])
elif mo == 3:
by, bx = divmod(b, nbx)
body += src_idx[by*4:by*4+4, bx*4:bx*4+4].tobytes()
return bytes(body)
def _idx(v):
"""codebook index: 1 byte if it fits, else big-endian u16.
k>256 means 2-byte indices -- decided once by the header, not per block."""
v = int(v)
return bytes([v]) if _IDX_BYTES == 1 else struct.pack(">H", v)
_IDX_BYTES = 1
def main():
global _IDX_BYTES
ap = argparse.ArgumentParser()
ap.add_argument("frames_dir"); ap.add_argument("out")
ap.add_argument("--profile", choices=list(RC.PROFILES), default="sasi")
ap.add_argument("--lam", type=float, default=None)
ap.add_argument("--fps", type=int, default=12)
ap.add_argument("--iters", type=int, default=16)
ap.add_argument("--preview")
a = ap.parse_args()
prof = RC.PROFILES[a.profile]
lam = a.lam if a.lam is not None else prof["lam"]
k1, k4 = prof["k1"], prof["k4"]
_IDX_BYTES = 1 if max(k1, k4) <= 256 else 2
print(f"profile {a.profile}: {prof['desc']}")
print(f" target {prof['kbps']} KB/s, lam={lam}, k1={k1} k4={k4}, "
f"{_IDX_BYTES}-byte indices")
m = H.build(a.frames_dir, k1=k1, k4=k4, iters=a.iters)
enc = H.encode(m, lam=lam)
r = H.evaluate(m, enc, fps=a.fps)
H_, W_ = m["H"], m["W"]; nbx = W_ // 4
pal, idx = m["pal"], m["idx"]
# re-derive the per-frame symbols the same way encode() did
frames = []
for f, im in enumerate(idx):
B1 = H.blocks_of(im, pal, 4, 4); l1 = VQ.assign(B1, m["C1s"])
B4 = H.blocks_of(im, pal, 2, 2); l4 = VQ.assign(B4, m["C4s"])
q = H._group_2x2_into_4x4(np.arange(len(l4)), W_)
l4g = l4[q].reshape(-1, 4)
mode = enc["modes"][f]
frames.append(pack_modes(mode) + frame_payload(mode, l1, l4g, im, nbx))
palette = m["pal"][:256]
if len(palette) < 256:
palette = np.vstack([palette, np.zeros((256 - len(palette), 3), np.uint8)])
pal_b = palette.astype(np.uint8).tobytes()
cb1_b = m["cb1"].astype(np.uint8).tobytes()
cb4_b = m["cb4"].astype(np.uint8).tobytes()
off_pal = 32
off_cb1 = off_pal + len(pal_b)
off_cb4 = off_cb1 + len(cb1_b)
off_frm = off_cb4 + len(cb4_b)
hdr = (b"DLX1" + struct.pack(">HHHHHH", W_, H_, a.fps, len(idx), k1, k4)
+ struct.pack(">IIII", off_pal, off_cb1, off_cb4, off_frm))
assert len(hdr) == 32, len(hdr)
with open(a.out, "wb") as fh:
fh.write(hdr); fh.write(pal_b); fh.write(cb1_b); fh.write(cb4_b)
for p in frames:
fh.write(struct.pack(">I", len(p))); fh.write(p)
total = os.path.getsize(a.out)
vid = sum(len(p) + 4 for p in frames)
print(f" wrote {a.out}: {total} B "
f"(header+tables {total-vid} B, video {vid} B)")
print(f" {vid/len(idx):.0f} B/frame -> {vid/len(idx)*a.fps/1024:.1f} KB/s video"
f" + {RC.AUDIO_KBPS} KB/s audio = {vid/len(idx)*a.fps/1024+RC.AUDIO_KBPS:.1f} KB/s")
print(f" PSNR {r['psnr']:.2f} dB palette ceiling {r['pal']:.2f} dB "
f"loss {r['loss']:.2f} dB")
print(f" modes: SKIP {r['skip']:.1f}% V1 {r['v1']:.1f}% "
f"V4 {r['v4']:.1f}% RAW {r['raw']:.1f}%")
if a.preview:
from PIL import Image
f = len(idx) // 2
gap = np.full((H_ * 3, 4, 3), 40, np.uint8)
st = np.concatenate([VQ.zoom(m["rgb"][f], 3), gap,
VQ.zoom(pal[idx[f]], 3), gap,
VQ.zoom(pal[enc["recon"][f]], 3)], axis=1)
Image.fromarray(st).save(a.preview)
print(f" preview -> {a.preview} (source | palette ceiling | decoded)")
if __name__ == "__main__":
main()
+41
View File
@@ -0,0 +1,41 @@
#!/usr/bin/env python3
"""Extract decimated frames from a Blu-ray .m2ts into 256x192 PNGs.
Source is 1920x1080 (16:9). The arcade original is 4:3, so we CENTER-CROP to
1440x1080 by default -- see docs/STATUS.md open question on framing.
"""
import subprocess, sys, os, shutil
STREAM_DIR = "/media/reala-misaki/BDROM/BDMV/STREAM"
W, H = 256, 192
def duration(path):
out = subprocess.check_output(["ffprobe","-v","error","-show_entries",
"format=duration","-of","csv=p=0",path], text=True)
return float(out.strip())
def extract(stream, outdir, fps=12, mode="crop", start=None, dur=None):
src = f"{STREAM_DIR}/{stream}.m2ts"
total = duration(src)
if start is None: start = 0.0
if dur is None: dur = total - start
shutil.rmtree(outdir, ignore_errors=True); os.makedirs(outdir)
if mode == "crop": # 4:3 centre crop, arcade framing
vf = f"crop=1440:1080:240:0,hqdn3d=4:3:0:0,scale={W}:{H}:flags=lanczos"
elif mode == "squash": # full 16:9 squeezed into 4:3
vf = f"hqdn3d=4:3:0:0,scale={W}:{H}:flags=lanczos"
elif mode == "wide": # 16:9 preserved, letterboxed later
vf = f"hqdn3d=4:3:0:0,scale={W}:144:flags=lanczos"
else: raise ValueError(mode)
vf = f"fps={fps}," + vf
subprocess.check_call(["ffmpeg","-v","error","-ss",str(start),"-t",str(dur),
"-i",src,"-vf",vf,"-vsync","0",f"{outdir}/f%04d.png","-y"])
n = len(os.listdir(outdir))
print(f"{stream}: dur={total:.2f}s -> {n} frames @{fps}fps ({mode})")
return n
if __name__ == "__main__":
stream, outdir = sys.argv[1], sys.argv[2]
fps = int(sys.argv[3]) if len(sys.argv) > 3 else 12
mode = sys.argv[4] if len(sys.argv) > 4 else "crop"
extract(stream, outdir, fps, mode)
+99
View File
@@ -0,0 +1,99 @@
#!/usr/bin/env python3
"""Rate control: hit a target bitrate exactly, so one encoder serves both targets.
USER DECISION (session 2): ship TWO quality modes, SASI and SCSI. The codec's
bitrate ceiling is a build parameter; the encoder is otherwise identical.
Mechanism: the hybrid encoder's lagrangian `lam` trades distortion for bytes
monotonically, so per frame we binary-search lam to land inside a byte budget.
A leaky bucket lets a quiet frame bank bytes that an action frame can spend --
without that, quiet frames waste budget and action frames stay ugly.
The ceiling is HARD: the 68000 streams at a fixed rate off the disk, and a frame
that overruns is a dropped frame, not a slow frame.
"""
import numpy as np
import vq_hybrid as H
# Profiles. Bandwidths are the sustained-read figures the player can rely on;
# see docs/FINDINGS.md 5 -- these are FOLKLORE-grade until the disk benchmark
# is unblocked, so they are deliberately conservative fractions of the quoted
# ceiling (audio, seeks and container overhead come out of the same pipe).
# Calibrated against the CORRECTED rate-distortion measurement (FINDINGS 14).
#
# k=256 with 1-byte indices beats k=1024 with 2-byte indices at every matched
# bitrate. The earlier "+2.4 dB for k=1024" was an artifact of a rate model that
# charged 1 byte for a 10-bit index. 1-byte indices also mean the 68000 decoder
# reads a plain move.b with no alignment case, and the codebook is 8 KB not 32 KB.
#
# The two profiles are the SAME codec, decoder and bitstream -- only `lam` differs.
PROFILES = {
"sasi": dict(kbps=45, lam=300.0, k1=256, k4=256,
desc="stock 10MHz ACE/EXPERT, SASI",
quality="34.8 dB on 00020 / 28.3 dB on 00146"),
"scsi": dict(kbps=75, lam=100.0, k1=256, k4=256,
desc="Super/XVI, or CZ-6BS1 board in a 10MHz machine",
quality="35.9 dB on 00020 / 29.0 dB on 00146"),
}
# Not a shipping profile, but the curve continues: lam=25 is ~185 KB/s at ~38.7 dB
# with 26% RAW blocks, and lam->0 is pixel-exact (0.00 dB loss). Entropy-coding
# the payload (NOT YET IMPLEMENTED) should shift the whole curve ~1.4x left.
AUDIO_KBPS = 7.8 # MSM6258 ADPCM 15.6kHz mono -- comes out of the same budget
def frame_budget(kbps, fps=12, audio=AUDIO_KBPS):
"""bytes per video frame after audio takes its cut"""
return (kbps - audio) * 1024.0 / fps
def encode_rate_controlled(m, target_kbps, fps=12, bucket_frames=8,
lam_lo=1.0, lam_hi=2e5, steps=9, verbose=False):
budget = frame_budget(target_kbps, fps)
bucket = 0.0 # banked bytes, capped at bucket_frames*budget
cap = bucket_frames * budget
out_recon, out_modes, out_sizes, out_lam = [], [], [], []
# encode() is whole-sequence; drive it per-lam and pick per frame.
# Cheaper than re-running the whole encoder per frame: precompute the ladder.
ladder = []
lams = np.geomspace(lam_lo, lam_hi, steps)
for lam in lams:
e = H.encode(m, lam=float(lam))
ladder.append(e)
if verbose:
print(f" lam={lam:9.0f} mean {e['sizes'].mean():6.0f} B/frame")
nf = len(m["idx"])
for f in range(nf):
allow = budget + bucket
# cheapest lam (highest quality) whose size fits the allowance
pick = len(lams) - 1
for i in range(len(lams)):
if ladder[i]["sizes"][f] <= allow:
pick = i; break
sz = ladder[pick]["sizes"][f]
bucket = min(cap, bucket + budget - sz)
out_recon.append(ladder[pick]["recon"][f])
out_modes.append(ladder[pick]["modes"][f])
out_sizes.append(sz); out_lam.append(lams[pick])
return dict(recon=out_recon, modes=out_modes, sizes=np.array(out_sizes),
lam=np.array(out_lam), nb=ladder[0]["nb"], budget=budget)
def summarise(m, enc, target_kbps, fps=12):
import vq as VQ
pal = m["pal"]
rec = [pal[i] for i in enc["recon"]]
src = [pal[i] for i in m["idx"]]
p = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], rec)])
pp = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], src)])
sz = enc["sizes"]
mo = np.concatenate(enc["modes"])
return dict(target=target_kbps, psnr=p, pal=pp, loss=pp - p,
mean_B=sz.mean(), max_B=sz.max(), budget=enc["budget"],
kbps=sz.mean() * fps / 1024 + AUDIO_KBPS,
over=100.0 * np.mean(sz > enc["budget"]),
skip=100 * (mo == 0).mean(), v1=100 * (mo == 1).mean(),
v4=100 * (mo == 2).mean())
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env python3
"""4x4 vector-quantisation prototype for the X68000 Dragon's Lair port.
Pipeline mirrors what the 68000 decoder would actually do, so the measured
quality is honest:
frames -> per-scene 256-colour palette (median cut, NO dither)
-> 4x4 blocks of PALETTISED rgb
-> k-means codebook (luma-weighted euclidean)
-> each codeword's 16 pixels snapped back to a palette index
The decoder only ever copies 16 palette indices out of a table, so the codebook
entries MUST be legal palette indices -- both quantisation losses compose.
No sklearn on this box; k-means is hand-rolled (chunked, numpy).
"""
import numpy as np, glob, os, sys
from PIL import Image
BW = BH = 4 # block size
# ITU-R BT.601 luma weights, squared -- we compare in a luma-weighted RGB space
LUMA = np.array([0.299, 0.587, 0.114], dtype=np.float32)
def load_frames(d):
fs = sorted(glob.glob(f"{d}/f*.png"))
return [np.asarray(Image.open(f).convert("RGB")) for f in fs]
def scene_palette(rgb, colors=256, stride=3):
"""One shared palette for the whole scene, no dithering (cel art is flat)."""
samp = np.concatenate([r.reshape(-1, 3) for r in rgb[::stride]])
ref = Image.fromarray(samp.reshape(-1, 1, 3)).quantize(
colors=colors, method=Image.MEDIANCUT, dither=Image.NONE)
pal = np.array(ref.getpalette()[:colors * 3], dtype=np.uint8).reshape(-1, 3)
return ref, pal
def palettise(rgb, ref):
return [np.asarray(Image.fromarray(r).quantize(palette=ref, dither=Image.NONE),
dtype=np.uint8) for r in rgb]
def blockify(idx, pal, bw=BW, bh=BH):
"""(H,W) palette indices -> (nblocks, bh*bw*3) float32 luma-weighted RGB."""
H, W = idx.shape
rgb = pal[idx].astype(np.float32) * LUMA # weight once, up front
b = rgb.reshape(H // bh, bh, W // bw, bw, 3).transpose(0, 2, 1, 3, 4)
return b.reshape(-1, bh * bw * 3)
def kmeans(X, k, iters=24, seed=0):
"""Chunked Lloyd's algorithm. k-means++ style seeding, deterministic."""
rng = np.random.default_rng(seed)
n = X.shape[0]
if n <= k:
return X.copy(), np.arange(n)
# seed: farthest-point sampling on a random subsample (cheap k-means++)
sub = X[rng.choice(n, min(n, 20000), replace=False)]
C = np.empty((k, X.shape[1]), dtype=np.float32)
C[0] = sub[rng.integers(len(sub))]
d2 = ((sub - C[0]) ** 2).sum(1)
for i in range(1, k):
C[i] = sub[np.argmax(d2)]
d2 = np.minimum(d2, ((sub - C[i]) ** 2).sum(1))
lab = None
for _ in range(iters):
lab = assign(X, C)
newC = C.copy()
cnt = np.bincount(lab, minlength=k)
s = np.zeros_like(C)
np.add.at(s, lab, X)
nz = cnt > 0
newC[nz] = s[nz] / cnt[nz, None]
# revive dead codewords on the worst-fit blocks
if (~nz).any():
err = ((X - newC[lab]) ** 2).sum(1)
worst = np.argsort(err)[-int((~nz).sum()):]
newC[~nz] = X[worst]
if np.allclose(newC, C):
C = newC; break
C = newC
return C, assign(X, C)
def assign(X, C, chunk=8192):
"""Nearest centroid, chunked to bound memory."""
Cn = (C ** 2).sum(1)
out = np.empty(X.shape[0], dtype=np.int32)
for i in range(0, X.shape[0], chunk):
x = X[i:i + chunk]
d = Cn[None, :] - 2.0 * (x @ C.T) # + |x|^2, constant per row
out[i:i + chunk] = np.argmin(d, axis=1)
return out
def snap_codebook(C, pal, bw=BW, bh=BH):
"""Centroids (luma-weighted RGB) -> legal palette indices, as the ROM stores them."""
cb_rgb = C.reshape(-1, bh * bw, 3) / LUMA # undo the weighting
palw = pal.astype(np.float32) * LUMA
flat = (cb_rgb * LUMA).reshape(-1, 3)
d = (flat ** 2).sum(1)[:, None] - 2 * (flat @ palw.T) + (palw ** 2).sum(1)[None, :]
return np.argmin(d, axis=1).astype(np.uint8).reshape(-1, bh * bw)
def unblockify(labels, cb_idx, H, W, bw=BW, bh=BH):
blocks = cb_idx[labels].reshape(H // bh, W // bw, bh, bw)
return blocks.transpose(0, 2, 1, 3).reshape(H, W)
def psnr(a, b):
mse = np.mean((a.astype(np.float64) - b.astype(np.float64)) ** 2)
return 99.0 if mse == 0 else 10 * np.log10(255.0 ** 2 / mse)
def encode_scene(frames_dir, k=256, bw=BW, bh=BH, iters=24):
rgb = load_frames(frames_dir)
H, W = rgb[0].shape[:2]
ref, pal = scene_palette(rgb)
idx = palettise(rgb, ref)
X = np.concatenate([blockify(i, pal, bw, bh) for i in idx])
C, _ = kmeans(X, k, iters)
cb_idx = snap_codebook(C, pal, bw, bh)
# re-assign against the SNAPPED codebook: that's what the decoder can produce
Csnap = (pal[cb_idx].astype(np.float32) * LUMA).reshape(k, -1)
recon, labels = [], []
for i in idx:
lab = assign(blockify(i, pal, bw, bh), Csnap)
labels.append(lab)
recon.append(unblockify(lab, cb_idx, H, W, bw, bh))
return dict(rgb=rgb, pal=pal, idx=idx, cb_idx=cb_idx, labels=labels,
recon=recon, H=H, W=W, k=k, bw=bw, bh=bh)
def report(r, name=""):
pal, idx, recon = r["pal"], r["idx"], r["recon"]
src8 = [pal[i] for i in idx]
vq8 = [pal[i] for i in recon]
orig = r["rgb"]
p_pal = np.mean([psnr(o, s) for o, s in zip(orig, src8)])
p_vq = np.mean([psnr(o, v) for o, v in zip(orig, vq8)])
p_vq_only = np.mean([psnr(s, v) for s, v in zip(src8, vq8)])
nb = (r["H"] // r["bh"]) * (r["W"] // r["bw"])
idxbits = int(np.ceil(np.log2(r["k"])))
keyf = nb * idxbits / 8
# block-delta cost: how many block indices change frame to frame
ch = [np.count_nonzero(r["labels"][i] != r["labels"][i - 1]) / nb
for i in range(1, len(r["labels"]))]
print(f"--- {name} k={r['k']} block={r['bw']}x{r['bh']} ---")
print(f" palette-only PSNR : {p_pal:5.2f} dB (floor: 256c is the best we can do)")
print(f" after VQ PSNR : {p_vq:5.2f} dB (loss from VQ alone: {p_pal-p_vq:.2f} dB)")
print(f" VQ vs palettised : {p_vq_only:5.2f} dB")
print(f" blocks/frame : {nb} keyframe {keyf:.0f} B codebook {r['k']*r['bw']*r['bh']} B")
if ch:
print(f" blocks changed/frm: mean {100*np.mean(ch):5.1f}% p90 {100*np.percentile(ch,90):5.1f}%")
return dict(p_pal=p_pal, p_vq=p_vq, nb=nb, keyf=keyf,
chg=np.mean(ch) if ch else 0, chg90=np.percentile(ch,90) if ch else 0)
def zoom(a, f=3):
return np.repeat(np.repeat(a, f, axis=0), f, axis=1)
def compare_png(r, frame, out, f=3):
pal = r["pal"]
src = pal[r["idx"][frame]]
vq = pal[r["recon"][frame]]
orig = r["rgb"][frame]
gap = np.full((src.shape[0] * f, 4, 3), 40, dtype=np.uint8)
strip = np.concatenate([zoom(orig, f), gap, zoom(src, f), gap, zoom(vq, f)], axis=1)
Image.fromarray(strip).save(out)
return out
+150
View File
@@ -0,0 +1,150 @@
#!/usr/bin/env python3
"""Cinepak-style hybrid VQ with a RAW escape: per-4x4-block choice of
SKIP / V1 (one 4x4 codeword) / V4 (four 2x2 codewords) / RAW (16 literal indices).
Flat 4x4 VQ at k=256 visibly destroys Bluth's ink linework (see docs/FINDINGS.md).
The standard fix is to let detailed blocks spend 4x the bits. Rate control picks
the split per block by rate-distortion, so the bitrate ceiling stays deterministic
-- which is the whole reason we chose VQ over a lossless delta.
The RAW mode is what makes ONE codec serve both shipping targets (session 2
user decision: SASI and SCSI quality modes). As lam -> 0 the encoder buys RAW
blocks until the frame is pixel-exact against the palettised source, so the
SCSI profile is not a second codec -- it is the same bitstream with the rate
knob opened up. The 68000 decoder needs no extra path: RAW is a straight copy,
which is cheaper than V4.
Bitstream per frame (what the 68000 actually parses):
2 bits/block header, packed: 00=SKIP 01=V1 10=V4 11=RAW
then the payload in block order: V1 -> 1 index, V4 -> 4, RAW -> 16
"""
import numpy as np, sys
from PIL import Image
import vq as VQ
LUMA = VQ.LUMA
def blocks_of(idx, pal, bw, bh):
return VQ.blockify(idx, pal, bw, bh)
def build(frames_dir, k1=256, k4=256, iters=16, lam=0.0):
rgb = VQ.load_frames(frames_dir)
H, W = rgb[0].shape[:2]
ref, pal = VQ.scene_palette(rgb)
idx = VQ.palettise(rgb, ref)
# --- two codebooks, trained on the whole scene ---
X1 = np.concatenate([blocks_of(i, pal, 4, 4) for i in idx])
C1, _ = VQ.kmeans(X1, k1, iters)
cb1 = VQ.snap_codebook(C1, pal, 4, 4) # (k1,16) palette idx
C1s = (pal[cb1].astype(np.float32) * LUMA).reshape(k1, -1)
X4 = np.concatenate([blocks_of(i, pal, 2, 2) for i in idx])
C4, _ = VQ.kmeans(X4, k4, iters)
cb4 = VQ.snap_codebook(C4, pal, 2, 2) # (k4,4) palette idx
C4s = (pal[cb4].astype(np.float32) * LUMA).reshape(k4, -1)
return dict(rgb=rgb, pal=pal, idx=idx, H=H, W=W,
cb1=cb1, C1s=C1s, cb4=cb4, C4s=C4s, k1=k1, k4=k4)
def _v1_recon(lab1, cb1, H, W):
return VQ.unblockify(lab1, cb1, H, W, 4, 4)
def encode(m, lam=0.02, skip_thresh=0.0, idx_bytes=None):
"""lam = lagrangian rate weight (bytes -> squared-error units).
Higher lam => more V1/SKIP => smaller & softer.
idx_bytes: size of ONE codebook index in the bitstream. k>256 needs 2 bytes,
which doubles what V1 and V4 actually cost -- if the RD model ignores that
it systematically over-picks V4 and under-reports the bitrate. Defaults to
the value implied by the codebook sizes."""
if idx_bytes is None:
idx_bytes = 1 if max(m["k1"], m["k4"]) <= 256 else 2
pal, idx, H, W = m["pal"], m["idx"], m["H"], m["W"]
nbx, nby = W // 4, H // 4
nb = nbx * nby
recon, modes, sizes = [], [], []
prev = None
for f, im in enumerate(idx):
B1 = blocks_of(im, pal, 4, 4) # (nb,48)
l1 = VQ.assign(B1, m["C1s"])
e1 = ((B1 - m["C1s"][l1]) ** 2).sum(1)
B4 = blocks_of(im, pal, 2, 2) # (nb*4,12) in 2x2 raster
l4 = VQ.assign(B4, m["C4s"])
e4raw = ((B4 - m["C4s"][l4]) ** 2).sum(1)
# regroup 2x2 blocks (raster over 8x12... ) into their parent 4x4 block
q = _group_2x2_into_4x4(np.arange(nb * 4), W)
e4 = e4raw[q].reshape(nb, 4).sum(1)
l4g = l4[q].reshape(nb, 4)
# SKIP: cost of reusing the previous *reconstructed* block
if prev is None:
eS = np.full(nb, np.inf)
else:
pb = blocks_of(prev, pal, 4, 4)
eS = ((B1 - pb) ** 2).sum(1)
# RAW: zero distortion against the palettised source, 16 bytes
eR = np.zeros(nb)
# rate-distortion choice: true byte cost per mode. The 2-bit header is
# paid by every block regardless, so it drops out of the comparison.
bV1 = 1.0 * idx_bytes
bV4 = 4.0 * idx_bytes
bRAW = 16.0 # RAW is literal palette bytes, never indices
cost = np.stack([eS + lam * 0.0, e1 + lam * bV1,
e4 + lam * bV4, eR + lam * bRAW])
mode = np.argmin(cost, axis=0).astype(np.uint8)
out = np.empty((H, W), dtype=np.uint8)
_paint(out, mode, l1, l4g, m["cb1"], m["cb4"], prev, nbx, nby, im)
recon.append(out); modes.append(mode)
nV1 = int((mode == 1).sum()); nV4 = int((mode == 2).sum())
nR = int((mode == 3).sum())
sizes.append(nb * 2 / 8 + (nV1 + nV4 * 4) * idx_bytes + nR * 16)
prev = out
return dict(recon=recon, modes=modes, sizes=np.array(sizes), nb=nb)
def _group_2x2_into_4x4(a, W):
"""map 2x2-block raster order -> (nb4, 4) grouping by parent 4x4 block"""
n2x = W // 2
n2y = len(a) // n2x
g = a.reshape(n2y, n2x)
g = g.reshape(n2y // 2, 2, n2x // 2, 2).transpose(0, 2, 1, 3)
return g.reshape(-1)
def _paint(out, mode, l1, l4g, cb1, cb4, prev, nbx, nby, src):
for b in range(len(mode)):
by, bx = divmod(b, nbx)
y, x = by * 4, bx * 4
mo = mode[b]
if mo == 0:
out[y:y+4, x:x+4] = prev[y:y+4, x:x+4]
elif mo == 1:
out[y:y+4, x:x+4] = cb1[l1[b]].reshape(4, 4)
elif mo == 3:
out[y:y+4, x:x+4] = src[y:y+4, x:x+4]
else:
c = cb4[l4g[b]].reshape(2, 2, 2, 2) # (sub_y,sub_x,2,2)
out[y:y+2, x:x+2] = c[0, 0]; out[y:y+2, x+2:x+4] = c[0, 1]
out[y+2:y+4, x:x+2] = c[1, 0]; out[y+2:y+4, x+2:x+4] = c[1, 1]
def evaluate(m, enc, fps=12):
pal = m["pal"]
rec = [pal[i] for i in enc["recon"]]
src = [pal[i] for i in m["idx"]]
p_vq = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], rec)])
p_pal = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], src)])
mo = np.concatenate(enc["modes"])
sz = enc["sizes"].mean()
return dict(psnr=p_vq, pal=p_pal, loss=p_pal - p_vq, bytes=sz,
kbps=sz * fps / 1024,
skip=100 * (mo == 0).mean(), v1=100 * (mo == 1).mean(),
v4=100 * (mo == 2).mean(), raw=100 * (mo == 3).mean())