Session 2: hybrid VQ codec, two quality profiles, three corrections
Answers session 1's critical-path question. Flat 4x4 VQ at k=256 was prototyped and REJECTED by eye: Dirk's face disintegrates and ink outlines break into 4-pixel stair-steps. The 256-colour palettised frame is excellent, so the palette was never the problem -- block VQ was. Replaced it with a Cinepak-style hybrid: each 4x4 block is SKIP, one 4x4 codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion. The RAW escape makes lam=0 pixel-exact (measured 0.00 dB loss), so the quality knob spans lossless to heavily-compressed in one bitstream. Per the user's decision, ships TWO quality profiles from that one codec, one decoder and one bitstream -- only the rate knob differs: sasi 45 KB/s lam=300 34.8 dB stock 10MHz ACE/EXPERT scsi 75 KB/s lam=100 35.9 dB Super/XVI or CZ-6BS1 Three corrections to earlier numbers: 1. Session 1's "183 KB/s at 12fps" was a bad extrapolation. Halving the framerate does not halve the bitrate -- decimation roughly doubles the per-frame delta. Re-measured directly: 340 KB/s for session 1's own RLE, 247 KB/s for changed-spans+deflate. The lossless floor is 319 MB. 2. A FOURTH false-good result, same family as the three in FINDINGS 4: k=1024 codebooks appeared to buy +2.4 dB free, because the rate model charged 1 byte for a 10-bit index. Charging the true cost reverses the verdict -- k=256 wins at every matched bitrate, and by 5 dB at the low end where the SASI profile lives. k=256 ships. 3. Stream inventory: the ~3-5MB clips are 1.2-1.7s, not ~60s, and some 60s streams are menus, not content. Any survey must classify before averaging. Also cleared both candidate sources for the game-logic layer: the SNES project is MIT and DirkSimple is zlib, so the arcade scene graph can be imported and the two transcriptions diffed against each other. Encoder is working end-to-end: extract.py -> vq/vq_hybrid/ratectl -> encode.py, emitting a big-endian DLX1 container the 68000 can parse with plain moves. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -8,3 +8,6 @@ assets/audio/
|
|||||||
assets/frames/
|
assets/frames/
|
||||||
build/
|
build/
|
||||||
roms/
|
roms/
|
||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
|
*.dlx
|
||||||
|
|||||||
@@ -19,9 +19,24 @@ docs/ findings, status, hardware reference
|
|||||||
tools/analysis/ frame-analysis scripts (01/02 marked BROKEN as regression refs)
|
tools/analysis/ frame-analysis scripts (01/02 marked BROKEN as regression refs)
|
||||||
tools/bench/ MAME Lua injection harness + 68000 benchmark sources
|
tools/bench/ MAME Lua injection harness + 68000 benchmark sources
|
||||||
tools/vasm/ vasm m68k assembler (built from source)
|
tools/vasm/ vasm m68k assembler (built from source)
|
||||||
tools/encoder/ VQ encoder (not yet written)
|
tools/encoder/ hybrid VQ encoder + DLX1 container writer (working)
|
||||||
src/player/ 68000 player (not yet written)
|
src/player/ 68000 player (not yet written)
|
||||||
assets/ extracted frames/audio (gitignored)
|
assets/ extracted frames/audio (gitignored)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Encoder
|
||||||
|
|
||||||
|
```
|
||||||
|
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
|
||||||
|
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile sasi --preview p.png
|
||||||
|
```
|
||||||
|
|
||||||
|
Two quality profiles ship from one codec and one decoder — `sasi` (45 KB/s) and
|
||||||
|
`scsi` (120 KB/s) are two points on the same rate-distortion curve. The codec is
|
||||||
|
a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four
|
||||||
|
2x2 codewords, or RAW literal pixels, chosen per block by rate-distortion.
|
||||||
|
|
||||||
|
The RAW escape means `lam=0` is pixel-exact against the palettised frame, so the
|
||||||
|
quality knob spans lossless to heavily-compressed without changing the bitstream.
|
||||||
|
|
||||||
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored — supply your own.
|
Source media (`DRAGONS_LAIR.iso`) and ROMs are gitignored — supply your own.
|
||||||
|
|||||||
@@ -188,3 +188,185 @@ Their 516 chapters are finer-grained than our 224 Blu-ray streams, so mapping
|
|||||||
their event table onto our footage means subdividing streams by timecode.
|
their event table onto our footage means subdividing streams by timecode.
|
||||||
|
|
||||||
Caveat: all of the above is from README/repo-tree summaries, not their source.
|
Caveat: all of the above is from README/repo-tree summaries, not their source.
|
||||||
|
|
||||||
|
---
|
||||||
|
---
|
||||||
|
|
||||||
|
# Findings — session 2 (2026-08-23)
|
||||||
|
|
||||||
|
## 8. CORRECTION to session 1: halving the framerate does NOT halve the bitrate
|
||||||
|
|
||||||
|
Session 1 measured 365 KB/s for naive delta+RLE at 24 fps and wrote
|
||||||
|
"(~183 KB/s at 12fps)". **That extrapolation is wrong.** Decimating to 12 fps
|
||||||
|
roughly doubles the per-frame delta, so the *rate* stays nearly flat.
|
||||||
|
|
||||||
|
Re-measured directly on 12 fps decimated frames (4 scenes, 66 frames):
|
||||||
|
|
||||||
|
| codec (all LOSSLESS w.r.t. the 256-colour frame) | B/frame | KB/s @12 | 22 min | ratio |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| raw 8bpp 256x192 | 49152 | 576 | 743 MB | 1.0:1 |
|
||||||
|
| session 1 row-span + RLE | 29055 | 340 | 439 MB | 1.7:1 |
|
||||||
|
| XOR vs prev + deflate | 30196 | 354 | 456 MB | 1.6:1 |
|
||||||
|
| **changed-spans + deflate** | **21110** | **247** | **319 MB** | **2.3:1** |
|
||||||
|
| changed-spans + LZMA | 18759 | 220 | 283 MB | 2.6:1 |
|
||||||
|
|
||||||
|
Session 1's own RLE re-measured at 12 fps gives **340 KB/s, not 183**.
|
||||||
|
Any plan that assumed 183 KB/s was based on a bad number.
|
||||||
|
|
||||||
|
Deflate-class entropy coding on top of the span payload is worth **1.4x** over
|
||||||
|
hand-rolled RLE, and LZ decode is cheap on a 68000 (byte copies), so the
|
||||||
|
lossless floor is ~247 KB/s / 319 MB. That is **infeasible on SASI** and
|
||||||
|
**tight but real on SCSI**.
|
||||||
|
|
||||||
|
## 9. Flat 4x4 VQ at k=256 is NOT acceptable — confirmed by eye
|
||||||
|
|
||||||
|
The risk flagged in 6 is real. At k=256, 4x4:
|
||||||
|
|
||||||
|
| scene | palette-only PSNR | after VQ | VQ loss |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 00010 | 38.35 | 29.68 | 8.67 dB |
|
||||||
|
| 00020 | 39.90 | 32.67 | 7.22 dB |
|
||||||
|
| 00146 | 35.25 | 29.35 | 5.89 dB |
|
||||||
|
| 00181 | 41.92 | 32.87 | 9.05 dB |
|
||||||
|
|
||||||
|
Visually: Dirk's face disintegrates, teeth and eyes turn to mush, ink outlines
|
||||||
|
break into 4-pixel stair-steps, colour bleeds across block boundaries.
|
||||||
|
|
||||||
|

|
||||||
|
*Left: 1080p source. Middle: 256-colour palettised 256x192 — the quality ceiling,
|
||||||
|
and it is excellent. Right: flat 4x4 VQ at k=256. This is the result that killed
|
||||||
|
the flat-VQ architecture.*
|
||||||
|
|
||||||
|
**Crucially, the 256-colour palettised frame itself looks excellent.** Flat cel
|
||||||
|
art with a per-scene median-cut palette and no dithering is near-transparent
|
||||||
|
(35-42 dB). So the palette is not the problem and 256 colours is not the
|
||||||
|
problem — **block VQ is**. The quality ceiling we should hold ourselves to is
|
||||||
|
the palettised frame, not the 1080p source.
|
||||||
|
|
||||||
|
## 10. Hybrid VQ (Cinepak V1/V4 + SKIP) — this is the codec
|
||||||
|
|
||||||
|
Per 4x4 block, choose by rate-distortion: SKIP (reuse previous frame),
|
||||||
|
V1 (one 4x4 codeword, 1 byte), or V4 (four 2x2 codewords, 4 bytes),
|
||||||
|
with a 2-bit-per-block mode header. `lam` is the lagrangian rate knob.
|
||||||
|
|
||||||
|
Measured, k1=k4=256, 4 scenes (mean of the per-scene table in the session log):
|
||||||
|
|
||||||
|
| lam | PSNR | loss vs palette | SKIP% | V1% | V4% | B/frame | KB/s @12 |
|
||||||
|
|---|---|---|---|---|---|---|---|
|
||||||
|
| 0 (max quality) | 33.9 | 4.9 | 30.8 | 18.5 | 50.8 | 7574 | 88.8 |
|
||||||
|
| 200 | 31.9 | 5.9 | 44.0 | 37.6 | 18.4 | 4183 | 49.0 |
|
||||||
|
| 1000 | 31.6 | 7.3 | 47.4 | 47.7 | 4.9 | 2841 | 33.3 |
|
||||||
|
| 5000 | 25.5 | 13.3 | 55.6 | 44.4 | 0.0 | 2134 | 25.0 |
|
||||||
|
|
||||||
|
At a **matched ~30 KB/s** the hybrid beats flat 4x4 VQ by ~1 dB, and unlike flat
|
||||||
|
VQ it keeps scaling: at 89 KB/s it reaches within **4.9 dB of the palette
|
||||||
|
ceiling**, which flat VQ cannot reach at any bitrate.
|
||||||
|
|
||||||
|
Note V4% collapses to 0 at lam=5000 — that is the knob doing exactly what it
|
||||||
|
should: under a hard ceiling, detail blocks are the first thing sacrificed.
|
||||||
|
|
||||||
|
## 11. Codebook size sweep (flat 4x4, for reference)
|
||||||
|
|
||||||
|
| block | k | PSNR | loss | key B | changed% | KB/s @12 | codebook RAM |
|
||||||
|
|---|---|---|---|---|---|---|---|
|
||||||
|
| 4x4 | 256 | 30.46 | 8.39 | 3072 | 52.7 | 28.5 | 8K |
|
||||||
|
| 4x4 | 1024 | 32.89 | 5.96 | 3840 | 56.6 | 35.6 | 32K |
|
||||||
|
|
||||||
|
+2.4 dB for 24K more RAM and 7 KB/s. With 2 MB of RAM, a 1024-entry codebook is
|
||||||
|
cheap and clearly worth it. (RAM figure is the word-expanded form the blitter
|
||||||
|
wants: k * 16 px * 2 bytes.)
|
||||||
|
|
||||||
|
## 12. Source framing — OPEN
|
||||||
|
|
||||||
|
The Blu-ray is **full-frame 1920x1080 16:9 with no pillarboxing**. The arcade
|
||||||
|
original is 4:3. The extractor currently centre-crops 1440x1080, which is the
|
||||||
|
arcade-faithful choice but discards image the 2006 remaster added. Options are
|
||||||
|
`crop` (default), `squash`, `wide` in `tools/encoder/extract.py`.
|
||||||
|
**Not yet decided; needs an eyeball comparison against arcade reference.**
|
||||||
|
|
||||||
|
## 13. Stream inventory correction
|
||||||
|
|
||||||
|
Session 1 said "typical scene clip ~60s". Sampled directly: the ~3-5 MB streams
|
||||||
|
are **1.2-1.7 s** clips — these are the individual arcade death/action moments,
|
||||||
|
which is exactly the granularity the game logic needs. Some 60 s streams
|
||||||
|
(e.g. 00203) are **menu screens, not content**. Any survey must classify
|
||||||
|
menu vs content before averaging, or the bitrate numbers are diluted by static
|
||||||
|
menus.
|
||||||
|
|
||||||
|
## 14. A FOURTH false-good result — and the correction
|
||||||
|
|
||||||
|
Add this to the 4 list. The mechanism was new but the shape was identical.
|
||||||
|
|
||||||
|
**The false result:** flat and hybrid VQ both showed **+2.4 dB for k=1024 over
|
||||||
|
k=256** at an apparently similar bitrate, which made a 1024-entry codebook look
|
||||||
|
like an obvious win. The k=1024 quality ladder rendered from that run looked
|
||||||
|
great at "45 KB/s".
|
||||||
|
|
||||||
|
**The bug:** the rate-distortion model in `vq_hybrid.encode()` charged **1 byte**
|
||||||
|
per codebook index unconditionally. A 1024-entry codebook needs a **10-bit index,
|
||||||
|
stored as 2 bytes**. So every k=1024 measurement understated the V1 and V4
|
||||||
|
payload by exactly 2x, *and* the lagrangian mode decision was choosing V4 on the
|
||||||
|
belief that four codewords cost 4 bytes when they cost 8.
|
||||||
|
|
||||||
|
**After charging the true index cost** (`idx_bytes` is now explicit and defaults
|
||||||
|
from the codebook size), matched-bitrate comparison on scene 00020:
|
||||||
|
|
||||||
|
| KB/s | k=256 (1-byte idx) | k=1024 (2-byte idx) |
|
||||||
|
|---|---|---|
|
||||||
|
| ~32-42 | **33.87 dB** @ 32.5 | 28.91 dB @ 42.3 |
|
||||||
|
| ~44-52 | **34.80 dB** @ 44.1 | 35.13 dB @ 52.5 |
|
||||||
|
| ~72-86 | **35.87 dB** @ 72.2 | 36.51 dB @ 86.0 |
|
||||||
|
|
||||||
|
k=1024 buys +0.3 to +0.6 dB for +19% bitrate — a wash at best — and at the low
|
||||||
|
end where the SASI profile lives it is **5 dB worse**, because the 2-byte index
|
||||||
|
floor dominates once V4 is priced out.
|
||||||
|
|
||||||
|
**k=256 with 1-byte indices is the shipping choice.** It is also the better
|
||||||
|
decoder: a plain `move.b` index with no alignment case, and an 8 KB codebook
|
||||||
|
instead of 32 KB.
|
||||||
|
|
||||||
|
**The general lesson, again:** the comparison was not wrong about VQ, it was
|
||||||
|
wrong about *cost*. When a knob looks like a free win, check that the rate model
|
||||||
|
is charging for it. Same failure family as 4.1-4.3: a plausible number produced
|
||||||
|
by a pipeline that was not measuring what it claimed to measure.
|
||||||
|
|
||||||
|
## 15. Rate-distortion curve of the shipping codec (k=256, corrected)
|
||||||
|
|
||||||
|
Scene 00020 (Dirk screaming, close-up face — the hardest case for linework),
|
||||||
|
and 00146. Includes the 2-bit-per-block mode header. No entropy coding yet.
|
||||||
|
|
||||||
|
| lam | 00020 PSNR | 00020 KB/s | 00146 PSNR | 00146 KB/s | SKIP | V1 | V4 | RAW |
|
||||||
|
|---|---|---|---|---|---|---|---|---|
|
||||||
|
| 25 | 38.68 | 182.2 | 31.04 | 193.5 | ~37% | ~24% | ~13% | ~26% |
|
||||||
|
| 100 | 35.87 | 72.2 | 29.04 | 72.5 | ~41% | ~34% | ~21% | ~4% |
|
||||||
|
| 300 | 34.80 | 44.1 | 28.28 | 44.4 | ~44% | ~42% | ~14% | 0% |
|
||||||
|
| 800 | 33.87 | 32.5 | 27.77 | 36.1 | ~46% | ~48% | ~5% | 0% |
|
||||||
|
| 2000 | 27.57 | 25.5 | 24.88 | 30.2 | ~50% | ~49% | ~1% | 0% |
|
||||||
|
|
||||||
|
Palette ceilings: 00020 = 39.90 dB, 00146 = 35.25 dB.
|
||||||
|
|
||||||
|

|
||||||
|
*The shipping codec across the rate knob. Top: source, palette ceiling, lam=25.
|
||||||
|
Bottom: lam=100 (`scsi` profile), lam=300 (`sasi` profile), lam=800.
|
||||||
|
Both shipping profiles hold Bluth's linework; the failure only starts past lam=800.*
|
||||||
|
|
||||||
|
Two things to read off this table:
|
||||||
|
- **The cliff is between lam=800 and lam=2000.** That is where V4 is priced out
|
||||||
|
entirely and detail blocks have nowhere to go. Do not ship past lam~800.
|
||||||
|
- **RAW is doing real work at high bitrate** (26% of blocks at lam=25) and
|
||||||
|
vanishes by lam=300. It is what makes the top of the curve reach the palette
|
||||||
|
ceiling, and it costs the decoder nothing — RAW is the cheapest mode to blit.
|
||||||
|
|
||||||
|
## 16. Licences cleared for the game-logic layer
|
||||||
|
|
||||||
|
Both checked this session:
|
||||||
|
|
||||||
|
- **astrobleem/SNES-SuperDragonsLairArcade — MIT**, "Copyright (c) 2026 Chad
|
||||||
|
Doebelin". `data/events/` holds 516 XML chapter definitions with timing and
|
||||||
|
event data. Reusable with attribution.
|
||||||
|
- **icculus/DirkSimple — zlib.** Independent from-scratch reimplementation of
|
||||||
|
the game logic in Lua, scene/timing tables in `game.lua`. Also permissive.
|
||||||
|
|
||||||
|
Having **two independent permissively-licensed transcriptions** of the arcade
|
||||||
|
scene graph is better than one: they can be diffed against each other to catch
|
||||||
|
transcription errors before any of it is committed to 68000 tables.
|
||||||
|
|||||||
+104
-103
@@ -1,4 +1,4 @@
|
|||||||
# Status & next-session handoff — end of session 1 (2026-08-23)
|
# Status & next-session handoff — end of session 2 (2026-08-23)
|
||||||
|
|
||||||
## Decisions locked
|
## Decisions locked
|
||||||
|
|
||||||
@@ -7,133 +7,134 @@
|
|||||||
| Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint |
|
| Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint |
|
||||||
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
|
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
|
||||||
| Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway |
|
| Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway |
|
||||||
| Codec | 4x4 vector quantization, per-scene codebook + block delta | CPU is idle, I/O is the ceiling — spend cycles to buy bandwidth |
|
| **Codec** | **hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion** | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 |
|
||||||
|
| **Quality modes** | **two: `sasi` and `scsi`** (USER DECISION, session 2) | one codec, one decoder, one bitstream; only `lam` differs |
|
||||||
| Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win |
|
| Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win |
|
||||||
| Medium | SCSI HDD image (.hds) | but see SASI/SCSI split below |
|
|
||||||
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
|
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
|
||||||
|
| SNES project reuse | **MIT — cleared** | `data/events/` scene graph is reusable with attribution |
|
||||||
|
|
||||||
**OPEN QUESTION for the user:** stock 10MHz machines are **SASI**, not SCSI.
|
### The SASI/SCSI question is RESOLVED
|
||||||
Three options, not yet chosen:
|
Session 1 left "which machine do we target" open. The user's answer: **ship both**,
|
||||||
1. Stock 10MHz + SASI (purist) — VQ becomes mandatory
|
as two quality profiles. This is now implemented rather than hypothetical — the
|
||||||
2. Stock 10MHz + CZ-6BS1 SCSI board — relieves I/O, keeps CPU honest
|
bitrate ceiling is a build parameter in `tools/encoder/ratectl.py`:
|
||||||
3. Super/XVI baseline — built-in SCSI, still a 10MHz 68000
|
|
||||||
|
|
||||||
Recommendation: make the codec's bitrate ceiling a **build parameter**, so one
|
| profile | target | lam | quality (00020 / 00146) | machine |
|
||||||
encoder serves all three and the target is chosen at package time.
|
|---|---|---|---|---|
|
||||||
|
| `sasi` | 45 KB/s | 300 | 34.8 / 28.3 dB | stock 10MHz ACE/EXPERT |
|
||||||
|
| `scsi` | 75 KB/s | 100 | 35.9 / 29.0 dB | Super/XVI, or CZ-6BS1 board |
|
||||||
|
|
||||||
|
Codebooks are **k=256 with 1-byte indices** in both profiles. k=1024 was measured
|
||||||
|
and rejected — see FINDINGS 14, it was a false-good result from a rate model
|
||||||
|
that undercharged the index. Do not ship past `lam~800`; FINDINGS 15 has the cliff.
|
||||||
|
|
||||||
|
Because of the RAW escape mode, `lam=0` is **pixel-exact** against the palettised
|
||||||
|
frame (measured 0.00 dB loss). The profiles are two points on one continuous
|
||||||
|
rate-distortion curve, not two codecs.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Working setup
|
## What session 2 settled
|
||||||
|
|
||||||
**MAME ROMs** — `~/mame/roms/x68000.zip` (present, working).
|
1. **The critical-path question is answered.** "Does VQ soften Bluth's linework
|
||||||
Must pass **`-bios ipl10`**; the default BIOS is `cz600ce`, whose split
|
unacceptably?" — **flat 4x4 k=256 VQ: yes, badly. Hybrid VQ with k=1024: no.**
|
||||||
even/odd IPL halves (`rh-ix0897cezz.ic12` / `rh-ix0898cezz.ic11`) are absent.
|
Verified by eye, not just PSNR. See `docs/FINDINGS.md` 9-11.
|
||||||
`-verifyroms` will still report those two as missing — this is expected and harmless.
|
2. **Session 1's 12fps bitrate was wrong** (183 KB/s claimed, 340 KB/s measured).
|
||||||
|
Halving the framerate does not halve the bitrate. FINDINGS 8.
|
||||||
|
2b. **A fourth false-good result was produced and caught this session** — k=1024
|
||||||
|
codebooks looked like a +2.4 dB free win because the rate model charged 1 byte
|
||||||
|
for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
|
||||||
|
3. **The 256-colour palettised frame is the real quality ceiling** and it looks
|
||||||
|
excellent. Judge the codec against that, not against 1080p.
|
||||||
|
4. Encoder exists and produces a real bitstream: `tools/encoder/`.
|
||||||
|
|
||||||
Boots headless at ~430-480% speed:
|
---
|
||||||
|
|
||||||
|
## Encoder — working
|
||||||
|
|
||||||
|
```
|
||||||
|
python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
|
||||||
|
python3 tools/encoder/encode.py /tmp/fr_00020 out.dlx --profile sasi --preview p.png
|
||||||
|
```
|
||||||
|
|
||||||
|
| file | role |
|
||||||
|
|---|---|
|
||||||
|
| `extract.py` | .m2ts -> 256x192 PNGs, 12fps, spatial-only denoise |
|
||||||
|
| `vq.py` | palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR |
|
||||||
|
| `vq_hybrid.py` | the codec: 4 block modes + lagrangian mode decision |
|
||||||
|
| `ratectl.py` | SASI/SCSI profiles, leaky-bucket rate control |
|
||||||
|
| `encode.py` | CLI + `DLX1` container writer |
|
||||||
|
|
||||||
|
`DLX1` container layout is documented in the `encode.py` docstring. All
|
||||||
|
multi-byte fields are **big-endian** so the 68000 reads them with a plain `move`.
|
||||||
|
|
||||||
|
### Known encoder gaps
|
||||||
|
- **Rate control is written but not yet wired into `encode.py`** — the CLI uses a
|
||||||
|
fixed `lam` from the profile. `ratectl.encode_rate_controlled()` exists and
|
||||||
|
builds a lam-ladder per frame; it needs hooking up and validating.
|
||||||
|
- **Payload is not entropy-coded.** Deflate on the payload should buy ~1.4x
|
||||||
|
(measured on the lossless path, FINDINGS 8). LZ decode is cheap on a 68000.
|
||||||
|
- Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
|
||||||
|
- `_paint` is a Python per-block loop — fine for prototyping, slow for a full
|
||||||
|
disc encode. Vectorise before the 224-stream run.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Working setup (unchanged from session 1, re-verified)
|
||||||
|
|
||||||
|
**MAME ROMs** — `~/mame/roms/x68000.zip`. Must pass **`-bios ipl10`**.
|
||||||
```
|
```
|
||||||
mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3
|
mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3
|
||||||
```
|
```
|
||||||
|
**Assembler** — `tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s`
|
||||||
|
|
||||||
**Assembler** — vasm built from source, binary at `tools/vasm/vasmm68k_mot`
|
**Blu-ray** — `udisksctl loop-setup -r -f DRAGONS_LAIR.iso` -> `/media/reala-misaki/BDROM`
|
||||||
(source tarball alongside it). Verified correct 68000 output.
|
(still mounted as of end of session 2).
|
||||||
```
|
|
||||||
tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s
|
|
||||||
```
|
|
||||||
|
|
||||||
**Blu-ray** — mount with:
|
**MAME Lua harness** — `tools/bench/*.lua`, working. Three gotchas (retain the
|
||||||
```
|
notifier subscription in a global; the stack register is `SP` not `A7`;
|
||||||
udisksctl loop-setup -r -f DRAGONS_LAIR.iso # -> /media/reala-misaki/BDROM
|
`autoboot_script` fires at PC=0 before boot) are documented in FINDINGS.
|
||||||
```
|
|
||||||
NOTE: this loop mount is still active from session 1. Re-mount if the machine rebooted.
|
**Two shell traps, both hit again this session:**
|
||||||
|
- piping MAME (or any long job) through `grep` block-buffers — write to a file.
|
||||||
|
- `pkill -f <pattern>` matches your own shell and kills it (exit 144).
|
||||||
|
Use `pkill -x` or kill by PID.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## MAME Lua harness — WORKING, reusable
|
## STILL BLOCKED: disk throughput benchmark
|
||||||
|
|
||||||
`tools/bench/*.lua` inject 68000 machine code straight into emulated RAM and time
|
Unchanged from session 1 — `IOCS _B_READ` returns -1 uniformly. Full diagnosis
|
||||||
it against the emulated clock. No bootable disk or OS required. This is the
|
and the four untested hypotheses are in session 1's notes (git history of this
|
||||||
measurement rig for all future cycle-cost work (blit timing, decoder benchmarks).
|
file, commit 65112b9).
|
||||||
|
|
||||||
Pattern:
|
**This now matters more than session 1 thought.** Session 1 dismissed it because
|
||||||
```
|
"VQ at 30 KB/s is correct whether SASI does 300 or 600 KB/s". But we now ship
|
||||||
mame x68000 -bios ipl10 -video none -sound none -nothrottle \
|
*two profiles*, and the profile bitrates (45 / 120 KB/s) are set against
|
||||||
-seconds_to_run 30 -plugins -autoboot_script yourscript.lua
|
**folklore** bandwidth figures. A real measurement would let us set them
|
||||||
```
|
honestly instead of conservatively. Next move is the untried SCSI path:
|
||||||
|
`-exp1 cz6bs1 -hard disk.chd`.
|
||||||
### Three MAME Lua gotchas — all cost real time, all now solved
|
|
||||||
1. **Retain the notifier subscription.** `emu.add_machine_frame_notifier()` returns
|
|
||||||
a token; if you drop it into a chunk-local it is garbage-collected and the
|
|
||||||
callback **silently stops firing**. Assign it to a **global** (`SUB = ...`).
|
|
||||||
2. **The stack pointer is `SP`, not `A7`** in `cpu.state[...]`.
|
|
||||||
Full list: A0-A6, D0-D7, PC, SP, SR, USP, CURPC, CURFLAGS, IR.
|
|
||||||
3. **`autoboot_script` fires at time=0, before boot** (PC=0). Wait until
|
|
||||||
`machine.time` >= ~5s before injecting, or IOCS is not yet initialised.
|
|
||||||
|
|
||||||
Also: piping MAME through `grep` block-buffers output — write raw to a file when
|
|
||||||
backgrounding, or you will see an empty log and assume a hang.
|
|
||||||
And never `pkill -f 'mame x68000'` — the pattern matches your own shell and kills it
|
|
||||||
(exit 144). Use `pkill -x mame`.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## BLOCKED: disk throughput benchmark
|
|
||||||
|
|
||||||
**Goal:** measure real SASI/SCSI KB/s to replace the folklore figures in FINDINGS.md §5.
|
|
||||||
|
|
||||||
**Status:** harness fully working; the IOCS call itself fails.
|
|
||||||
|
|
||||||
`IOCS _B_READ` ($46 via `TRAP #15`; d1.hb=PDA, d2.l=position, d3.l=bytes, a1=buffer)
|
|
||||||
returns **`FFFFFFFF` (-1), zero reads**, uniformly across:
|
|
||||||
- all 16 PDA values $80-$8F
|
|
||||||
- both d1 encodings (PDA in bits 31-24 and bits 15-8)
|
|
||||||
- image sizes 10MB / 20MB / 40MB
|
|
||||||
|
|
||||||
The uniformity is the diagnostic: calls are **dispatched and cleanly rejected**,
|
|
||||||
so `TRAP #15` and IOCS are reachable. MAME does mount the image
|
|
||||||
(`:x68k_hdc: opened image file bench.hdf`).
|
|
||||||
|
|
||||||
**Untested hypotheses, in rough order of likelihood:**
|
|
||||||
1. The raw image has no X68000 SASI format, so the IPL's boot scan never registered
|
|
||||||
a usable drive and IOCS refuses. Would need Human68k to format one — **we have
|
|
||||||
no Human68k image on this system.**
|
|
||||||
2. MAME's `x68k_hdc` SASI implementation may be too partial for IOCS-level reads.
|
|
||||||
3. `SP=$8000` may put the injected stack on top of the IOCS work area in low RAM.
|
|
||||||
Try a much higher stack.
|
|
||||||
4. The **SCSI path was never tried** — this is the obvious next move and is more
|
|
||||||
relevant to the target anyway:
|
|
||||||
`-exp1 cz6bs1 -hard disk.chd` with `exp1:cz6bs1:scsi:0 harddisk`
|
|
||||||
(`-listmedia` gains a `harddisk` slot accepting .chd/.hd/.hdv/.2mg/.hdi).
|
|
||||||
|
|
||||||
**Honest assessment: this benchmark is NOT on the critical path.** The VQ codec
|
|
||||||
(~30 KB/s) is correct whether SASI does 300 or 600 KB/s. Do not let it block the
|
|
||||||
encoder. Its real value is deciding whether the *simpler* row-span codec could
|
|
||||||
have sufficed.
|
|
||||||
|
|
||||||
Caveat if resumed: MAME idealizes drive seek latency. That's acceptable because the
|
|
||||||
realistic deployment is BlueSCSI/SCSI2SD (SD-backed, no mechanical seek), so what
|
|
||||||
gets measured is the bus/DMAC/controller path — the genuine ceiling. The caveat
|
|
||||||
only bites for a real period spinning drive.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Next steps, in priority order
|
## Next steps, in priority order
|
||||||
|
|
||||||
1. **Build the VQ encoder** (`tools/encoder/`) — 4x4 blocks, per-scene codebook,
|
1. **Wire rate control into `encode.py`** and validate that the hard ceiling
|
||||||
block delta. Emit sample PNGs for visual evaluation. **The open question is
|
actually holds on an action scene (the whole point of choosing VQ).
|
||||||
whether VQ softens Bluth's ink linework unacceptably — decide by eye before
|
2. **Entropy-code the payload** (deflate) — ~1.4x for cheap 68000 decode cost.
|
||||||
committing to the architecture.**
|
3. **68000 decoder skeleton**: parse `DLX1`, expand codebooks to word-per-pixel,
|
||||||
2. **Full-disc survey** — all 224 streams, not 5s samples, to firm up bitrate
|
blit V1/V4/RAW/SKIP. Measure real cycles with the existing MAME Lua harness —
|
||||||
(current numbers are +/-30%) and map streams onto the arcade scene graph.
|
this is the first time the harness gets used for its actual purpose.
|
||||||
3. **Check the SNES project's license**, then evaluate reusing `data/events/`
|
4. **Full-disc survey** — classify menu vs content first (FINDINGS 13), then
|
||||||
(516 chapters / 29 scenes) as the scene-graph and input-timing layer.
|
measure bitrate across all 224 streams per profile.
|
||||||
4. Resolve the SASI/SCSI target question with the user.
|
5. **Resolve the framing question** (FINDINGS 12: crop vs squash vs wide).
|
||||||
5. Optionally unblock the disk benchmark via the SCSI path.
|
6. Unblock the disk benchmark via the SCSI path, then re-set profile bitrates.
|
||||||
6. 68000 player skeleton: CRTC init for 256x192x256c, `movem.l` blitter,
|
7. Import the SNES project's `data/events/` (MIT, cleared) as the scene graph.
|
||||||
ADPCM via HD63450 DMA.
|
Cross-check against DirkSimple (zlib) which has the same data independently.
|
||||||
|
8. ADPCM audio: MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in `ratectl`,
|
||||||
|
not yet extracted or encoded.
|
||||||
|
|
||||||
## Not yet started
|
## Not yet started
|
||||||
- Any 68000 player code
|
- Any 68000 player code
|
||||||
- ADPCM audio extraction/encoding (MSM6258, 15.6kHz mono, ~7.8 KB/s, ~10MB for 22min)
|
- ADPCM audio extraction/encoding
|
||||||
- Disk image packaging / container format
|
- Disk image packaging
|
||||||
- Game logic (scene branching, input windows, death clips)
|
- Game logic (scene branching, input windows, death clips)
|
||||||
|
|||||||
Binary file not shown.
|
After Width: | Height: | Size: 680 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 1.3 MiB |
@@ -0,0 +1,148 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Encode one scene to the DLX bitstream, at a chosen quality profile.
|
||||||
|
|
||||||
|
python3 tools/encoder/encode.py <frames_dir> <out.dlx> [--profile sasi|scsi]
|
||||||
|
[--lam N] [--fps 12] [--preview out.png]
|
||||||
|
|
||||||
|
Container (little-endian is WRONG here -- the 68000 is big-endian, so every
|
||||||
|
multi-byte field is big-endian and the decoder can read it with a plain move.w):
|
||||||
|
|
||||||
|
header, 32 bytes
|
||||||
|
0 'DLX1' magic
|
||||||
|
4 u16 width, u16 height
|
||||||
|
8 u16 fps, u16 nframes
|
||||||
|
12 u16 k1, u16 k4 codebook sizes
|
||||||
|
16 u32 palette offset (256 * 3 bytes, RGB888 -- the player converts
|
||||||
|
to the X68000's GRB555 at load time)
|
||||||
|
20 u32 cb1 offset (k1 * 16 bytes of palette indices)
|
||||||
|
24 u32 cb4 offset (k4 * 4 bytes)
|
||||||
|
28 u32 frames offset
|
||||||
|
then, per frame:
|
||||||
|
u32 payload length, then
|
||||||
|
ceil(nblocks*2/8) bytes of 2-bit mode headers, MSB-first, block raster order
|
||||||
|
then payloads in block order: V1 -> 1 byte, V4 -> 4 bytes, RAW -> 16 bytes
|
||||||
|
|
||||||
|
Codebooks are emitted as palette INDICES, not pixels. The player expands them
|
||||||
|
once at load time into word-per-pixel form so the blitter can movem them
|
||||||
|
straight into GVRAM -- k1=1024 costs 1024*16*2 = 32 KB of the 2 MB.
|
||||||
|
"""
|
||||||
|
import argparse, struct, sys, os
|
||||||
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
import numpy as np
|
||||||
|
import vq as VQ, vq_hybrid as H, ratectl as RC
|
||||||
|
|
||||||
|
|
||||||
|
def pack_modes(mode):
|
||||||
|
"""2 bits per block, MSB-first -- cheap for the 68000 to shift out."""
|
||||||
|
n = len(mode)
|
||||||
|
out = bytearray((n * 2 + 7) // 8)
|
||||||
|
for i, m in enumerate(mode):
|
||||||
|
out[i // 4] |= (int(m) & 3) << (6 - 2 * (i % 4))
|
||||||
|
return bytes(out)
|
||||||
|
|
||||||
|
|
||||||
|
def frame_payload(mode, l1, l4g, src_idx, nbx):
|
||||||
|
body = bytearray()
|
||||||
|
for b, mo in enumerate(mode):
|
||||||
|
if mo == 1:
|
||||||
|
body += _idx(l1[b])
|
||||||
|
elif mo == 2:
|
||||||
|
for j in range(4):
|
||||||
|
body += _idx(l4g[b][j])
|
||||||
|
elif mo == 3:
|
||||||
|
by, bx = divmod(b, nbx)
|
||||||
|
body += src_idx[by*4:by*4+4, bx*4:bx*4+4].tobytes()
|
||||||
|
return bytes(body)
|
||||||
|
|
||||||
|
|
||||||
|
def _idx(v):
|
||||||
|
"""codebook index: 1 byte if it fits, else big-endian u16.
|
||||||
|
k>256 means 2-byte indices -- decided once by the header, not per block."""
|
||||||
|
v = int(v)
|
||||||
|
return bytes([v]) if _IDX_BYTES == 1 else struct.pack(">H", v)
|
||||||
|
|
||||||
|
|
||||||
|
_IDX_BYTES = 1
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
global _IDX_BYTES
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("frames_dir"); ap.add_argument("out")
|
||||||
|
ap.add_argument("--profile", choices=list(RC.PROFILES), default="sasi")
|
||||||
|
ap.add_argument("--lam", type=float, default=None)
|
||||||
|
ap.add_argument("--fps", type=int, default=12)
|
||||||
|
ap.add_argument("--iters", type=int, default=16)
|
||||||
|
ap.add_argument("--preview")
|
||||||
|
a = ap.parse_args()
|
||||||
|
|
||||||
|
prof = RC.PROFILES[a.profile]
|
||||||
|
lam = a.lam if a.lam is not None else prof["lam"]
|
||||||
|
k1, k4 = prof["k1"], prof["k4"]
|
||||||
|
_IDX_BYTES = 1 if max(k1, k4) <= 256 else 2
|
||||||
|
|
||||||
|
print(f"profile {a.profile}: {prof['desc']}")
|
||||||
|
print(f" target {prof['kbps']} KB/s, lam={lam}, k1={k1} k4={k4}, "
|
||||||
|
f"{_IDX_BYTES}-byte indices")
|
||||||
|
|
||||||
|
m = H.build(a.frames_dir, k1=k1, k4=k4, iters=a.iters)
|
||||||
|
enc = H.encode(m, lam=lam)
|
||||||
|
r = H.evaluate(m, enc, fps=a.fps)
|
||||||
|
|
||||||
|
H_, W_ = m["H"], m["W"]; nbx = W_ // 4
|
||||||
|
pal, idx = m["pal"], m["idx"]
|
||||||
|
|
||||||
|
# re-derive the per-frame symbols the same way encode() did
|
||||||
|
frames = []
|
||||||
|
for f, im in enumerate(idx):
|
||||||
|
B1 = H.blocks_of(im, pal, 4, 4); l1 = VQ.assign(B1, m["C1s"])
|
||||||
|
B4 = H.blocks_of(im, pal, 2, 2); l4 = VQ.assign(B4, m["C4s"])
|
||||||
|
q = H._group_2x2_into_4x4(np.arange(len(l4)), W_)
|
||||||
|
l4g = l4[q].reshape(-1, 4)
|
||||||
|
mode = enc["modes"][f]
|
||||||
|
frames.append(pack_modes(mode) + frame_payload(mode, l1, l4g, im, nbx))
|
||||||
|
|
||||||
|
palette = m["pal"][:256]
|
||||||
|
if len(palette) < 256:
|
||||||
|
palette = np.vstack([palette, np.zeros((256 - len(palette), 3), np.uint8)])
|
||||||
|
pal_b = palette.astype(np.uint8).tobytes()
|
||||||
|
cb1_b = m["cb1"].astype(np.uint8).tobytes()
|
||||||
|
cb4_b = m["cb4"].astype(np.uint8).tobytes()
|
||||||
|
|
||||||
|
off_pal = 32
|
||||||
|
off_cb1 = off_pal + len(pal_b)
|
||||||
|
off_cb4 = off_cb1 + len(cb1_b)
|
||||||
|
off_frm = off_cb4 + len(cb4_b)
|
||||||
|
hdr = (b"DLX1" + struct.pack(">HHHHHH", W_, H_, a.fps, len(idx), k1, k4)
|
||||||
|
+ struct.pack(">IIII", off_pal, off_cb1, off_cb4, off_frm))
|
||||||
|
assert len(hdr) == 32, len(hdr)
|
||||||
|
|
||||||
|
with open(a.out, "wb") as fh:
|
||||||
|
fh.write(hdr); fh.write(pal_b); fh.write(cb1_b); fh.write(cb4_b)
|
||||||
|
for p in frames:
|
||||||
|
fh.write(struct.pack(">I", len(p))); fh.write(p)
|
||||||
|
|
||||||
|
total = os.path.getsize(a.out)
|
||||||
|
vid = sum(len(p) + 4 for p in frames)
|
||||||
|
print(f" wrote {a.out}: {total} B "
|
||||||
|
f"(header+tables {total-vid} B, video {vid} B)")
|
||||||
|
print(f" {vid/len(idx):.0f} B/frame -> {vid/len(idx)*a.fps/1024:.1f} KB/s video"
|
||||||
|
f" + {RC.AUDIO_KBPS} KB/s audio = {vid/len(idx)*a.fps/1024+RC.AUDIO_KBPS:.1f} KB/s")
|
||||||
|
print(f" PSNR {r['psnr']:.2f} dB palette ceiling {r['pal']:.2f} dB "
|
||||||
|
f"loss {r['loss']:.2f} dB")
|
||||||
|
print(f" modes: SKIP {r['skip']:.1f}% V1 {r['v1']:.1f}% "
|
||||||
|
f"V4 {r['v4']:.1f}% RAW {r['raw']:.1f}%")
|
||||||
|
|
||||||
|
if a.preview:
|
||||||
|
from PIL import Image
|
||||||
|
f = len(idx) // 2
|
||||||
|
gap = np.full((H_ * 3, 4, 3), 40, np.uint8)
|
||||||
|
st = np.concatenate([VQ.zoom(m["rgb"][f], 3), gap,
|
||||||
|
VQ.zoom(pal[idx[f]], 3), gap,
|
||||||
|
VQ.zoom(pal[enc["recon"][f]], 3)], axis=1)
|
||||||
|
Image.fromarray(st).save(a.preview)
|
||||||
|
print(f" preview -> {a.preview} (source | palette ceiling | decoded)")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Extract decimated frames from a Blu-ray .m2ts into 256x192 PNGs.
|
||||||
|
|
||||||
|
Source is 1920x1080 (16:9). The arcade original is 4:3, so we CENTER-CROP to
|
||||||
|
1440x1080 by default -- see docs/STATUS.md open question on framing.
|
||||||
|
"""
|
||||||
|
import subprocess, sys, os, shutil
|
||||||
|
|
||||||
|
STREAM_DIR = "/media/reala-misaki/BDROM/BDMV/STREAM"
|
||||||
|
W, H = 256, 192
|
||||||
|
|
||||||
|
def duration(path):
|
||||||
|
out = subprocess.check_output(["ffprobe","-v","error","-show_entries",
|
||||||
|
"format=duration","-of","csv=p=0",path], text=True)
|
||||||
|
return float(out.strip())
|
||||||
|
|
||||||
|
def extract(stream, outdir, fps=12, mode="crop", start=None, dur=None):
|
||||||
|
src = f"{STREAM_DIR}/{stream}.m2ts"
|
||||||
|
total = duration(src)
|
||||||
|
if start is None: start = 0.0
|
||||||
|
if dur is None: dur = total - start
|
||||||
|
shutil.rmtree(outdir, ignore_errors=True); os.makedirs(outdir)
|
||||||
|
if mode == "crop": # 4:3 centre crop, arcade framing
|
||||||
|
vf = f"crop=1440:1080:240:0,hqdn3d=4:3:0:0,scale={W}:{H}:flags=lanczos"
|
||||||
|
elif mode == "squash": # full 16:9 squeezed into 4:3
|
||||||
|
vf = f"hqdn3d=4:3:0:0,scale={W}:{H}:flags=lanczos"
|
||||||
|
elif mode == "wide": # 16:9 preserved, letterboxed later
|
||||||
|
vf = f"hqdn3d=4:3:0:0,scale={W}:144:flags=lanczos"
|
||||||
|
else: raise ValueError(mode)
|
||||||
|
vf = f"fps={fps}," + vf
|
||||||
|
subprocess.check_call(["ffmpeg","-v","error","-ss",str(start),"-t",str(dur),
|
||||||
|
"-i",src,"-vf",vf,"-vsync","0",f"{outdir}/f%04d.png","-y"])
|
||||||
|
n = len(os.listdir(outdir))
|
||||||
|
print(f"{stream}: dur={total:.2f}s -> {n} frames @{fps}fps ({mode})")
|
||||||
|
return n
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
stream, outdir = sys.argv[1], sys.argv[2]
|
||||||
|
fps = int(sys.argv[3]) if len(sys.argv) > 3 else 12
|
||||||
|
mode = sys.argv[4] if len(sys.argv) > 4 else "crop"
|
||||||
|
extract(stream, outdir, fps, mode)
|
||||||
@@ -0,0 +1,99 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Rate control: hit a target bitrate exactly, so one encoder serves both targets.
|
||||||
|
|
||||||
|
USER DECISION (session 2): ship TWO quality modes, SASI and SCSI. The codec's
|
||||||
|
bitrate ceiling is a build parameter; the encoder is otherwise identical.
|
||||||
|
|
||||||
|
Mechanism: the hybrid encoder's lagrangian `lam` trades distortion for bytes
|
||||||
|
monotonically, so per frame we binary-search lam to land inside a byte budget.
|
||||||
|
A leaky bucket lets a quiet frame bank bytes that an action frame can spend --
|
||||||
|
without that, quiet frames waste budget and action frames stay ugly.
|
||||||
|
|
||||||
|
The ceiling is HARD: the 68000 streams at a fixed rate off the disk, and a frame
|
||||||
|
that overruns is a dropped frame, not a slow frame.
|
||||||
|
"""
|
||||||
|
import numpy as np
|
||||||
|
import vq_hybrid as H
|
||||||
|
|
||||||
|
# Profiles. Bandwidths are the sustained-read figures the player can rely on;
|
||||||
|
# see docs/FINDINGS.md 5 -- these are FOLKLORE-grade until the disk benchmark
|
||||||
|
# is unblocked, so they are deliberately conservative fractions of the quoted
|
||||||
|
# ceiling (audio, seeks and container overhead come out of the same pipe).
|
||||||
|
# Calibrated against the CORRECTED rate-distortion measurement (FINDINGS 14).
|
||||||
|
#
|
||||||
|
# k=256 with 1-byte indices beats k=1024 with 2-byte indices at every matched
|
||||||
|
# bitrate. The earlier "+2.4 dB for k=1024" was an artifact of a rate model that
|
||||||
|
# charged 1 byte for a 10-bit index. 1-byte indices also mean the 68000 decoder
|
||||||
|
# reads a plain move.b with no alignment case, and the codebook is 8 KB not 32 KB.
|
||||||
|
#
|
||||||
|
# The two profiles are the SAME codec, decoder and bitstream -- only `lam` differs.
|
||||||
|
PROFILES = {
|
||||||
|
"sasi": dict(kbps=45, lam=300.0, k1=256, k4=256,
|
||||||
|
desc="stock 10MHz ACE/EXPERT, SASI",
|
||||||
|
quality="34.8 dB on 00020 / 28.3 dB on 00146"),
|
||||||
|
"scsi": dict(kbps=75, lam=100.0, k1=256, k4=256,
|
||||||
|
desc="Super/XVI, or CZ-6BS1 board in a 10MHz machine",
|
||||||
|
quality="35.9 dB on 00020 / 29.0 dB on 00146"),
|
||||||
|
}
|
||||||
|
# Not a shipping profile, but the curve continues: lam=25 is ~185 KB/s at ~38.7 dB
|
||||||
|
# with 26% RAW blocks, and lam->0 is pixel-exact (0.00 dB loss). Entropy-coding
|
||||||
|
# the payload (NOT YET IMPLEMENTED) should shift the whole curve ~1.4x left.
|
||||||
|
|
||||||
|
AUDIO_KBPS = 7.8 # MSM6258 ADPCM 15.6kHz mono -- comes out of the same budget
|
||||||
|
|
||||||
|
|
||||||
|
def frame_budget(kbps, fps=12, audio=AUDIO_KBPS):
|
||||||
|
"""bytes per video frame after audio takes its cut"""
|
||||||
|
return (kbps - audio) * 1024.0 / fps
|
||||||
|
|
||||||
|
|
||||||
|
def encode_rate_controlled(m, target_kbps, fps=12, bucket_frames=8,
|
||||||
|
lam_lo=1.0, lam_hi=2e5, steps=9, verbose=False):
|
||||||
|
budget = frame_budget(target_kbps, fps)
|
||||||
|
bucket = 0.0 # banked bytes, capped at bucket_frames*budget
|
||||||
|
cap = bucket_frames * budget
|
||||||
|
out_recon, out_modes, out_sizes, out_lam = [], [], [], []
|
||||||
|
|
||||||
|
# encode() is whole-sequence; drive it per-lam and pick per frame.
|
||||||
|
# Cheaper than re-running the whole encoder per frame: precompute the ladder.
|
||||||
|
ladder = []
|
||||||
|
lams = np.geomspace(lam_lo, lam_hi, steps)
|
||||||
|
for lam in lams:
|
||||||
|
e = H.encode(m, lam=float(lam))
|
||||||
|
ladder.append(e)
|
||||||
|
if verbose:
|
||||||
|
print(f" lam={lam:9.0f} mean {e['sizes'].mean():6.0f} B/frame")
|
||||||
|
|
||||||
|
nf = len(m["idx"])
|
||||||
|
for f in range(nf):
|
||||||
|
allow = budget + bucket
|
||||||
|
# cheapest lam (highest quality) whose size fits the allowance
|
||||||
|
pick = len(lams) - 1
|
||||||
|
for i in range(len(lams)):
|
||||||
|
if ladder[i]["sizes"][f] <= allow:
|
||||||
|
pick = i; break
|
||||||
|
sz = ladder[pick]["sizes"][f]
|
||||||
|
bucket = min(cap, bucket + budget - sz)
|
||||||
|
out_recon.append(ladder[pick]["recon"][f])
|
||||||
|
out_modes.append(ladder[pick]["modes"][f])
|
||||||
|
out_sizes.append(sz); out_lam.append(lams[pick])
|
||||||
|
|
||||||
|
return dict(recon=out_recon, modes=out_modes, sizes=np.array(out_sizes),
|
||||||
|
lam=np.array(out_lam), nb=ladder[0]["nb"], budget=budget)
|
||||||
|
|
||||||
|
|
||||||
|
def summarise(m, enc, target_kbps, fps=12):
|
||||||
|
import vq as VQ
|
||||||
|
pal = m["pal"]
|
||||||
|
rec = [pal[i] for i in enc["recon"]]
|
||||||
|
src = [pal[i] for i in m["idx"]]
|
||||||
|
p = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], rec)])
|
||||||
|
pp = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], src)])
|
||||||
|
sz = enc["sizes"]
|
||||||
|
mo = np.concatenate(enc["modes"])
|
||||||
|
return dict(target=target_kbps, psnr=p, pal=pp, loss=pp - p,
|
||||||
|
mean_B=sz.mean(), max_B=sz.max(), budget=enc["budget"],
|
||||||
|
kbps=sz.mean() * fps / 1024 + AUDIO_KBPS,
|
||||||
|
over=100.0 * np.mean(sz > enc["budget"]),
|
||||||
|
skip=100 * (mo == 0).mean(), v1=100 * (mo == 1).mean(),
|
||||||
|
v4=100 * (mo == 2).mean())
|
||||||
@@ -0,0 +1,174 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""4x4 vector-quantisation prototype for the X68000 Dragon's Lair port.
|
||||||
|
|
||||||
|
Pipeline mirrors what the 68000 decoder would actually do, so the measured
|
||||||
|
quality is honest:
|
||||||
|
|
||||||
|
frames -> per-scene 256-colour palette (median cut, NO dither)
|
||||||
|
-> 4x4 blocks of PALETTISED rgb
|
||||||
|
-> k-means codebook (luma-weighted euclidean)
|
||||||
|
-> each codeword's 16 pixels snapped back to a palette index
|
||||||
|
|
||||||
|
The decoder only ever copies 16 palette indices out of a table, so the codebook
|
||||||
|
entries MUST be legal palette indices -- both quantisation losses compose.
|
||||||
|
|
||||||
|
No sklearn on this box; k-means is hand-rolled (chunked, numpy).
|
||||||
|
"""
|
||||||
|
import numpy as np, glob, os, sys
|
||||||
|
from PIL import Image
|
||||||
|
|
||||||
|
BW = BH = 4 # block size
|
||||||
|
# ITU-R BT.601 luma weights, squared -- we compare in a luma-weighted RGB space
|
||||||
|
LUMA = np.array([0.299, 0.587, 0.114], dtype=np.float32)
|
||||||
|
|
||||||
|
|
||||||
|
def load_frames(d):
|
||||||
|
fs = sorted(glob.glob(f"{d}/f*.png"))
|
||||||
|
return [np.asarray(Image.open(f).convert("RGB")) for f in fs]
|
||||||
|
|
||||||
|
|
||||||
|
def scene_palette(rgb, colors=256, stride=3):
|
||||||
|
"""One shared palette for the whole scene, no dithering (cel art is flat)."""
|
||||||
|
samp = np.concatenate([r.reshape(-1, 3) for r in rgb[::stride]])
|
||||||
|
ref = Image.fromarray(samp.reshape(-1, 1, 3)).quantize(
|
||||||
|
colors=colors, method=Image.MEDIANCUT, dither=Image.NONE)
|
||||||
|
pal = np.array(ref.getpalette()[:colors * 3], dtype=np.uint8).reshape(-1, 3)
|
||||||
|
return ref, pal
|
||||||
|
|
||||||
|
|
||||||
|
def palettise(rgb, ref):
|
||||||
|
return [np.asarray(Image.fromarray(r).quantize(palette=ref, dither=Image.NONE),
|
||||||
|
dtype=np.uint8) for r in rgb]
|
||||||
|
|
||||||
|
|
||||||
|
def blockify(idx, pal, bw=BW, bh=BH):
|
||||||
|
"""(H,W) palette indices -> (nblocks, bh*bw*3) float32 luma-weighted RGB."""
|
||||||
|
H, W = idx.shape
|
||||||
|
rgb = pal[idx].astype(np.float32) * LUMA # weight once, up front
|
||||||
|
b = rgb.reshape(H // bh, bh, W // bw, bw, 3).transpose(0, 2, 1, 3, 4)
|
||||||
|
return b.reshape(-1, bh * bw * 3)
|
||||||
|
|
||||||
|
|
||||||
|
def kmeans(X, k, iters=24, seed=0):
|
||||||
|
"""Chunked Lloyd's algorithm. k-means++ style seeding, deterministic."""
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
n = X.shape[0]
|
||||||
|
if n <= k:
|
||||||
|
return X.copy(), np.arange(n)
|
||||||
|
# seed: farthest-point sampling on a random subsample (cheap k-means++)
|
||||||
|
sub = X[rng.choice(n, min(n, 20000), replace=False)]
|
||||||
|
C = np.empty((k, X.shape[1]), dtype=np.float32)
|
||||||
|
C[0] = sub[rng.integers(len(sub))]
|
||||||
|
d2 = ((sub - C[0]) ** 2).sum(1)
|
||||||
|
for i in range(1, k):
|
||||||
|
C[i] = sub[np.argmax(d2)]
|
||||||
|
d2 = np.minimum(d2, ((sub - C[i]) ** 2).sum(1))
|
||||||
|
lab = None
|
||||||
|
for _ in range(iters):
|
||||||
|
lab = assign(X, C)
|
||||||
|
newC = C.copy()
|
||||||
|
cnt = np.bincount(lab, minlength=k)
|
||||||
|
s = np.zeros_like(C)
|
||||||
|
np.add.at(s, lab, X)
|
||||||
|
nz = cnt > 0
|
||||||
|
newC[nz] = s[nz] / cnt[nz, None]
|
||||||
|
# revive dead codewords on the worst-fit blocks
|
||||||
|
if (~nz).any():
|
||||||
|
err = ((X - newC[lab]) ** 2).sum(1)
|
||||||
|
worst = np.argsort(err)[-int((~nz).sum()):]
|
||||||
|
newC[~nz] = X[worst]
|
||||||
|
if np.allclose(newC, C):
|
||||||
|
C = newC; break
|
||||||
|
C = newC
|
||||||
|
return C, assign(X, C)
|
||||||
|
|
||||||
|
|
||||||
|
def assign(X, C, chunk=8192):
|
||||||
|
"""Nearest centroid, chunked to bound memory."""
|
||||||
|
Cn = (C ** 2).sum(1)
|
||||||
|
out = np.empty(X.shape[0], dtype=np.int32)
|
||||||
|
for i in range(0, X.shape[0], chunk):
|
||||||
|
x = X[i:i + chunk]
|
||||||
|
d = Cn[None, :] - 2.0 * (x @ C.T) # + |x|^2, constant per row
|
||||||
|
out[i:i + chunk] = np.argmin(d, axis=1)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def snap_codebook(C, pal, bw=BW, bh=BH):
|
||||||
|
"""Centroids (luma-weighted RGB) -> legal palette indices, as the ROM stores them."""
|
||||||
|
cb_rgb = C.reshape(-1, bh * bw, 3) / LUMA # undo the weighting
|
||||||
|
palw = pal.astype(np.float32) * LUMA
|
||||||
|
flat = (cb_rgb * LUMA).reshape(-1, 3)
|
||||||
|
d = (flat ** 2).sum(1)[:, None] - 2 * (flat @ palw.T) + (palw ** 2).sum(1)[None, :]
|
||||||
|
return np.argmin(d, axis=1).astype(np.uint8).reshape(-1, bh * bw)
|
||||||
|
|
||||||
|
|
||||||
|
def unblockify(labels, cb_idx, H, W, bw=BW, bh=BH):
|
||||||
|
blocks = cb_idx[labels].reshape(H // bh, W // bw, bh, bw)
|
||||||
|
return blocks.transpose(0, 2, 1, 3).reshape(H, W)
|
||||||
|
|
||||||
|
|
||||||
|
def psnr(a, b):
|
||||||
|
mse = np.mean((a.astype(np.float64) - b.astype(np.float64)) ** 2)
|
||||||
|
return 99.0 if mse == 0 else 10 * np.log10(255.0 ** 2 / mse)
|
||||||
|
|
||||||
|
|
||||||
|
def encode_scene(frames_dir, k=256, bw=BW, bh=BH, iters=24):
|
||||||
|
rgb = load_frames(frames_dir)
|
||||||
|
H, W = rgb[0].shape[:2]
|
||||||
|
ref, pal = scene_palette(rgb)
|
||||||
|
idx = palettise(rgb, ref)
|
||||||
|
|
||||||
|
X = np.concatenate([blockify(i, pal, bw, bh) for i in idx])
|
||||||
|
C, _ = kmeans(X, k, iters)
|
||||||
|
cb_idx = snap_codebook(C, pal, bw, bh)
|
||||||
|
|
||||||
|
# re-assign against the SNAPPED codebook: that's what the decoder can produce
|
||||||
|
Csnap = (pal[cb_idx].astype(np.float32) * LUMA).reshape(k, -1)
|
||||||
|
recon, labels = [], []
|
||||||
|
for i in idx:
|
||||||
|
lab = assign(blockify(i, pal, bw, bh), Csnap)
|
||||||
|
labels.append(lab)
|
||||||
|
recon.append(unblockify(lab, cb_idx, H, W, bw, bh))
|
||||||
|
return dict(rgb=rgb, pal=pal, idx=idx, cb_idx=cb_idx, labels=labels,
|
||||||
|
recon=recon, H=H, W=W, k=k, bw=bw, bh=bh)
|
||||||
|
|
||||||
|
|
||||||
|
def report(r, name=""):
|
||||||
|
pal, idx, recon = r["pal"], r["idx"], r["recon"]
|
||||||
|
src8 = [pal[i] for i in idx]
|
||||||
|
vq8 = [pal[i] for i in recon]
|
||||||
|
orig = r["rgb"]
|
||||||
|
p_pal = np.mean([psnr(o, s) for o, s in zip(orig, src8)])
|
||||||
|
p_vq = np.mean([psnr(o, v) for o, v in zip(orig, vq8)])
|
||||||
|
p_vq_only = np.mean([psnr(s, v) for s, v in zip(src8, vq8)])
|
||||||
|
nb = (r["H"] // r["bh"]) * (r["W"] // r["bw"])
|
||||||
|
idxbits = int(np.ceil(np.log2(r["k"])))
|
||||||
|
keyf = nb * idxbits / 8
|
||||||
|
# block-delta cost: how many block indices change frame to frame
|
||||||
|
ch = [np.count_nonzero(r["labels"][i] != r["labels"][i - 1]) / nb
|
||||||
|
for i in range(1, len(r["labels"]))]
|
||||||
|
print(f"--- {name} k={r['k']} block={r['bw']}x{r['bh']} ---")
|
||||||
|
print(f" palette-only PSNR : {p_pal:5.2f} dB (floor: 256c is the best we can do)")
|
||||||
|
print(f" after VQ PSNR : {p_vq:5.2f} dB (loss from VQ alone: {p_pal-p_vq:.2f} dB)")
|
||||||
|
print(f" VQ vs palettised : {p_vq_only:5.2f} dB")
|
||||||
|
print(f" blocks/frame : {nb} keyframe {keyf:.0f} B codebook {r['k']*r['bw']*r['bh']} B")
|
||||||
|
if ch:
|
||||||
|
print(f" blocks changed/frm: mean {100*np.mean(ch):5.1f}% p90 {100*np.percentile(ch,90):5.1f}%")
|
||||||
|
return dict(p_pal=p_pal, p_vq=p_vq, nb=nb, keyf=keyf,
|
||||||
|
chg=np.mean(ch) if ch else 0, chg90=np.percentile(ch,90) if ch else 0)
|
||||||
|
|
||||||
|
|
||||||
|
def zoom(a, f=3):
|
||||||
|
return np.repeat(np.repeat(a, f, axis=0), f, axis=1)
|
||||||
|
|
||||||
|
|
||||||
|
def compare_png(r, frame, out, f=3):
|
||||||
|
pal = r["pal"]
|
||||||
|
src = pal[r["idx"][frame]]
|
||||||
|
vq = pal[r["recon"][frame]]
|
||||||
|
orig = r["rgb"][frame]
|
||||||
|
gap = np.full((src.shape[0] * f, 4, 3), 40, dtype=np.uint8)
|
||||||
|
strip = np.concatenate([zoom(orig, f), gap, zoom(src, f), gap, zoom(vq, f)], axis=1)
|
||||||
|
Image.fromarray(strip).save(out)
|
||||||
|
return out
|
||||||
@@ -0,0 +1,150 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Cinepak-style hybrid VQ with a RAW escape: per-4x4-block choice of
|
||||||
|
SKIP / V1 (one 4x4 codeword) / V4 (four 2x2 codewords) / RAW (16 literal indices).
|
||||||
|
|
||||||
|
Flat 4x4 VQ at k=256 visibly destroys Bluth's ink linework (see docs/FINDINGS.md).
|
||||||
|
The standard fix is to let detailed blocks spend 4x the bits. Rate control picks
|
||||||
|
the split per block by rate-distortion, so the bitrate ceiling stays deterministic
|
||||||
|
-- which is the whole reason we chose VQ over a lossless delta.
|
||||||
|
|
||||||
|
The RAW mode is what makes ONE codec serve both shipping targets (session 2
|
||||||
|
user decision: SASI and SCSI quality modes). As lam -> 0 the encoder buys RAW
|
||||||
|
blocks until the frame is pixel-exact against the palettised source, so the
|
||||||
|
SCSI profile is not a second codec -- it is the same bitstream with the rate
|
||||||
|
knob opened up. The 68000 decoder needs no extra path: RAW is a straight copy,
|
||||||
|
which is cheaper than V4.
|
||||||
|
|
||||||
|
Bitstream per frame (what the 68000 actually parses):
|
||||||
|
2 bits/block header, packed: 00=SKIP 01=V1 10=V4 11=RAW
|
||||||
|
then the payload in block order: V1 -> 1 index, V4 -> 4, RAW -> 16
|
||||||
|
"""
|
||||||
|
import numpy as np, sys
|
||||||
|
from PIL import Image
|
||||||
|
import vq as VQ
|
||||||
|
|
||||||
|
LUMA = VQ.LUMA
|
||||||
|
|
||||||
|
|
||||||
|
def blocks_of(idx, pal, bw, bh):
|
||||||
|
return VQ.blockify(idx, pal, bw, bh)
|
||||||
|
|
||||||
|
|
||||||
|
def build(frames_dir, k1=256, k4=256, iters=16, lam=0.0):
|
||||||
|
rgb = VQ.load_frames(frames_dir)
|
||||||
|
H, W = rgb[0].shape[:2]
|
||||||
|
ref, pal = VQ.scene_palette(rgb)
|
||||||
|
idx = VQ.palettise(rgb, ref)
|
||||||
|
|
||||||
|
# --- two codebooks, trained on the whole scene ---
|
||||||
|
X1 = np.concatenate([blocks_of(i, pal, 4, 4) for i in idx])
|
||||||
|
C1, _ = VQ.kmeans(X1, k1, iters)
|
||||||
|
cb1 = VQ.snap_codebook(C1, pal, 4, 4) # (k1,16) palette idx
|
||||||
|
C1s = (pal[cb1].astype(np.float32) * LUMA).reshape(k1, -1)
|
||||||
|
|
||||||
|
X4 = np.concatenate([blocks_of(i, pal, 2, 2) for i in idx])
|
||||||
|
C4, _ = VQ.kmeans(X4, k4, iters)
|
||||||
|
cb4 = VQ.snap_codebook(C4, pal, 2, 2) # (k4,4) palette idx
|
||||||
|
C4s = (pal[cb4].astype(np.float32) * LUMA).reshape(k4, -1)
|
||||||
|
return dict(rgb=rgb, pal=pal, idx=idx, H=H, W=W,
|
||||||
|
cb1=cb1, C1s=C1s, cb4=cb4, C4s=C4s, k1=k1, k4=k4)
|
||||||
|
|
||||||
|
|
||||||
|
def _v1_recon(lab1, cb1, H, W):
|
||||||
|
return VQ.unblockify(lab1, cb1, H, W, 4, 4)
|
||||||
|
|
||||||
|
|
||||||
|
def encode(m, lam=0.02, skip_thresh=0.0, idx_bytes=None):
|
||||||
|
"""lam = lagrangian rate weight (bytes -> squared-error units).
|
||||||
|
Higher lam => more V1/SKIP => smaller & softer.
|
||||||
|
|
||||||
|
idx_bytes: size of ONE codebook index in the bitstream. k>256 needs 2 bytes,
|
||||||
|
which doubles what V1 and V4 actually cost -- if the RD model ignores that
|
||||||
|
it systematically over-picks V4 and under-reports the bitrate. Defaults to
|
||||||
|
the value implied by the codebook sizes."""
|
||||||
|
if idx_bytes is None:
|
||||||
|
idx_bytes = 1 if max(m["k1"], m["k4"]) <= 256 else 2
|
||||||
|
pal, idx, H, W = m["pal"], m["idx"], m["H"], m["W"]
|
||||||
|
nbx, nby = W // 4, H // 4
|
||||||
|
nb = nbx * nby
|
||||||
|
recon, modes, sizes = [], [], []
|
||||||
|
prev = None
|
||||||
|
for f, im in enumerate(idx):
|
||||||
|
B1 = blocks_of(im, pal, 4, 4) # (nb,48)
|
||||||
|
l1 = VQ.assign(B1, m["C1s"])
|
||||||
|
e1 = ((B1 - m["C1s"][l1]) ** 2).sum(1)
|
||||||
|
|
||||||
|
B4 = blocks_of(im, pal, 2, 2) # (nb*4,12) in 2x2 raster
|
||||||
|
l4 = VQ.assign(B4, m["C4s"])
|
||||||
|
e4raw = ((B4 - m["C4s"][l4]) ** 2).sum(1)
|
||||||
|
# regroup 2x2 blocks (raster over 8x12... ) into their parent 4x4 block
|
||||||
|
q = _group_2x2_into_4x4(np.arange(nb * 4), W)
|
||||||
|
e4 = e4raw[q].reshape(nb, 4).sum(1)
|
||||||
|
l4g = l4[q].reshape(nb, 4)
|
||||||
|
|
||||||
|
# SKIP: cost of reusing the previous *reconstructed* block
|
||||||
|
if prev is None:
|
||||||
|
eS = np.full(nb, np.inf)
|
||||||
|
else:
|
||||||
|
pb = blocks_of(prev, pal, 4, 4)
|
||||||
|
eS = ((B1 - pb) ** 2).sum(1)
|
||||||
|
|
||||||
|
# RAW: zero distortion against the palettised source, 16 bytes
|
||||||
|
eR = np.zeros(nb)
|
||||||
|
|
||||||
|
# rate-distortion choice: true byte cost per mode. The 2-bit header is
|
||||||
|
# paid by every block regardless, so it drops out of the comparison.
|
||||||
|
bV1 = 1.0 * idx_bytes
|
||||||
|
bV4 = 4.0 * idx_bytes
|
||||||
|
bRAW = 16.0 # RAW is literal palette bytes, never indices
|
||||||
|
cost = np.stack([eS + lam * 0.0, e1 + lam * bV1,
|
||||||
|
e4 + lam * bV4, eR + lam * bRAW])
|
||||||
|
mode = np.argmin(cost, axis=0).astype(np.uint8)
|
||||||
|
|
||||||
|
out = np.empty((H, W), dtype=np.uint8)
|
||||||
|
_paint(out, mode, l1, l4g, m["cb1"], m["cb4"], prev, nbx, nby, im)
|
||||||
|
recon.append(out); modes.append(mode)
|
||||||
|
nV1 = int((mode == 1).sum()); nV4 = int((mode == 2).sum())
|
||||||
|
nR = int((mode == 3).sum())
|
||||||
|
sizes.append(nb * 2 / 8 + (nV1 + nV4 * 4) * idx_bytes + nR * 16)
|
||||||
|
prev = out
|
||||||
|
return dict(recon=recon, modes=modes, sizes=np.array(sizes), nb=nb)
|
||||||
|
|
||||||
|
|
||||||
|
def _group_2x2_into_4x4(a, W):
|
||||||
|
"""map 2x2-block raster order -> (nb4, 4) grouping by parent 4x4 block"""
|
||||||
|
n2x = W // 2
|
||||||
|
n2y = len(a) // n2x
|
||||||
|
g = a.reshape(n2y, n2x)
|
||||||
|
g = g.reshape(n2y // 2, 2, n2x // 2, 2).transpose(0, 2, 1, 3)
|
||||||
|
return g.reshape(-1)
|
||||||
|
|
||||||
|
|
||||||
|
def _paint(out, mode, l1, l4g, cb1, cb4, prev, nbx, nby, src):
|
||||||
|
for b in range(len(mode)):
|
||||||
|
by, bx = divmod(b, nbx)
|
||||||
|
y, x = by * 4, bx * 4
|
||||||
|
mo = mode[b]
|
||||||
|
if mo == 0:
|
||||||
|
out[y:y+4, x:x+4] = prev[y:y+4, x:x+4]
|
||||||
|
elif mo == 1:
|
||||||
|
out[y:y+4, x:x+4] = cb1[l1[b]].reshape(4, 4)
|
||||||
|
elif mo == 3:
|
||||||
|
out[y:y+4, x:x+4] = src[y:y+4, x:x+4]
|
||||||
|
else:
|
||||||
|
c = cb4[l4g[b]].reshape(2, 2, 2, 2) # (sub_y,sub_x,2,2)
|
||||||
|
out[y:y+2, x:x+2] = c[0, 0]; out[y:y+2, x+2:x+4] = c[0, 1]
|
||||||
|
out[y+2:y+4, x:x+2] = c[1, 0]; out[y+2:y+4, x+2:x+4] = c[1, 1]
|
||||||
|
|
||||||
|
|
||||||
|
def evaluate(m, enc, fps=12):
|
||||||
|
pal = m["pal"]
|
||||||
|
rec = [pal[i] for i in enc["recon"]]
|
||||||
|
src = [pal[i] for i in m["idx"]]
|
||||||
|
p_vq = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], rec)])
|
||||||
|
p_pal = np.mean([VQ.psnr(o, v) for o, v in zip(m["rgb"], src)])
|
||||||
|
mo = np.concatenate(enc["modes"])
|
||||||
|
sz = enc["sizes"].mean()
|
||||||
|
return dict(psnr=p_vq, pal=p_pal, loss=p_pal - p_vq, bytes=sz,
|
||||||
|
kbps=sz * fps / 1024,
|
||||||
|
skip=100 * (mo == 0).mean(), v1=100 * (mo == 1).mean(),
|
||||||
|
v4=100 * (mo == 2).mean(), raw=100 * (mo == 3).mean())
|
||||||
Reference in New Issue
Block a user