# Status & next-session handoff — end of session 2 (2026-08-23) ## Decisions locked | decision | value | why | |---|---|---| | Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint | | Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c | | Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway | | **Codec** | **hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion** | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 | | **Quality modes** | **two: `sasi` and `scsi`** (USER DECISION, session 2) | one codec, one decoder, one bitstream; only `lam` differs | | Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win | | Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something | | SNES project reuse | **MIT — cleared** | `data/events/` scene graph is reusable with attribution | ### The SASI/SCSI question is RESOLVED Session 1 left "which machine do we target" open. The user's answer: **ship both**, as two quality profiles. This is now implemented rather than hypothetical — the bitrate ceiling is a build parameter in `tools/encoder/ratectl.py`: | profile | target | lam | quality (00020 / 00146) | machine | |---|---|---|---|---| | `sasi` | 110 KB/s | 60 | 36.9 / 29.6 dB | stock 10MHz ACE/EXPERT | | `scsi` | 280 KB/s | 10 | 39.4 / 32.3 dB | Super/XVI, or CZ-6BS1 board | Sized against the user's working figure of **4 Mbps = 488 KB/s sustained**, on SD-backed SCSI (BlueSCSI / SCSI2SD) — so that rate is a bus-limited **constant**, not an average over seek latency. **Both profiles fit with room.** Ring-buffer simulation on the real per-frame sizes gives **zero required prefill** for every scene at both profiles: the fill delivers 40.69 KB per frame time and only one measured frame (42.10 KB) exceeds that, recovered by the next. A 256 KB buffer carries ~1 s of stall tolerance, far more than an SD-backed seek needs. FINDINGS 21. An earlier warning here said `scsi` did not fit because a frame peaked at 96.4% of the pipe. That compared instantaneous demand to a sustained rate as if they had to match frame-by-frame; with a buffer the test is cumulative, and it passes. `scsi` is now within **0.5 dB of the palette ceiling** on 00020. These were initially set at 45 / 75 KB/s, which was 12% / 7% bus utilisation — read off the RD curve rather than derived from the hardware. See FINDINGS 17. Codebooks are **k=256 with 1-byte indices** in both profiles. k=1024 was measured and rejected — see FINDINGS 14, it was a false-good result from a rate model that undercharged the index. Do not ship past `lam~800`; FINDINGS 15 has the cliff. Because of the RAW escape mode, `lam=0` is **pixel-exact** against the palettised frame (measured 0.00 dB loss). The profiles are two points on one continuous rate-distortion curve, not two codecs. --- ## What session 2 settled 1. **The critical-path question is answered.** "Does VQ soften Bluth's linework unacceptably?" — **flat 4x4 VQ: yes, badly. The hybrid (SKIP/V1/V4/RAW): no.** Verified by eye, not just PSNR. See `docs/FINDINGS.md` 9-11 and the two images in `docs/images/`. Both profiles use **k=256**; see item 2b. 2. **Session 1's 12fps bitrate was wrong** (183 KB/s claimed, 340 KB/s measured). Halving the framerate does not halve the bitrate. FINDINGS 8. 2b. **A fourth false-good result was produced and caught this session** — k=1024 codebooks looked like a +2.4 dB free win because the rate model charged 1 byte for a 10-bit index. FINDINGS 14. The k=256 configuration ships. 3. **The 256-colour palettised frame is the real quality ceiling** and it looks excellent. Judge the codec against that, not against 1080p. 4. Encoder exists and produces a real bitstream: `tools/encoder/`. --- ## Encoder — working ``` python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop python3 tools/encoder/encode.py /tmp/fr_00020 out.dlx --profile sasi --preview p.png ``` | file | role | |---|---| | `extract.py` | .m2ts -> 256x192 PNGs, 12fps, spatial-only denoise | | `vq.py` | palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR | | `vq_hybrid.py` | the codec: 4 block modes + lagrangian mode decision | | `ratectl.py` | SASI/SCSI profiles, leaky-bucket rate control | | `encode.py` | CLI + `DLX1` container writer | `DLX1` container layout is documented in the `encode.py` docstring. All multi-byte fields are **big-endian** so the 68000 reads them with a plain `move`. ### Known encoder gaps - **Rate control is written but not yet wired into `encode.py`** — the CLI uses a fixed `lam` from the profile. `ratectl.encode_rate_controlled()` exists and builds a lam-ladder per frame; it needs hooking up and validating. - **Payload is deliberately NOT entropy-coded** — deflate decode does not fit in the 68000's frame budget (FINDINGS 17.2). Do not "optimise" this later. - Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse. - `_paint` is a Python per-block loop — fine for prototyping, slow for a full disc encode. Vectorise before the 224-stream run. --- ## Working setup (unchanged from session 1, re-verified) **MAME ROMs** — `~/mame/roms/x68000.zip`. Must pass **`-bios ipl10`**. ``` mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3 ``` **Assembler** — `tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s` **Blu-ray** — `udisksctl loop-setup -r -f DRAGONS_LAIR.iso` -> `/media/reala-misaki/BDROM` (still mounted as of end of session 2). **MAME Lua harness** — `tools/bench/*.lua`, working. Three gotchas (retain the notifier subscription in a global; the stack register is `SP` not `A7`; `autoboot_script` fires at PC=0 before boot) are documented in FINDINGS. **Two shell traps, both hit again this session:** - piping MAME (or any long job) through `grep` block-buffers — write to a file. - `pkill -f ` matches your own shell and kills it (exit 144). Use `pkill -x` or kill by PID. - **`until ! pgrep -f foo.py; do sleep; done` watcher loops never exit.** The watching shell's own command line contains the string `foo.py`, so `pgrep -f` matches the watcher itself and the loop spins forever. Session 2 left 11 of these wedged for over an hour. Wait on the PID (`while kill -0 $PID`) or on a sentinel file the job touches when it finishes -- never on a `-f` name match. - **`timeout N mame ...` does not kill MAME.** MAME catches SIGTERM and, with an autoboot script blocked waiting on a flag that never arrives, never reaches its shutdown path. `timeout` without `-k` then waits forever while MAME burns a full core at `-nothrottle`. Always `timeout -k 5 N`. --- ## Disk throughput benchmark — still blocked, no longer gating `IOCS _B_READ` returns -1 uniformly. Full diagnosis and the four untested hypotheses are in session 1's notes (git history of this file, commit 65112b9); the ordered plan for retrying is in **`docs/BENCHMARK.md`**. **Status changed twice this session — read this rather than the git history.** It was briefly promoted to critical-path while the working bandwidth figure was misread as 4 MB/s. With the correct figure (**4 Mbps = 488 KB/s**) and the ring-buffer simulation showing **zero required prefill** for both profiles (FINDINGS 21), the design no longer hangs on it. Pixel-exact on SCSI is **not** available at 4 Mbps — it needs 92-97% of the pipe — so there is no longer a "measure it and maybe ship transparent" decision waiting. What the benchmark is still worth doing for: - **Confirming the 4 Mbps figure.** It is user-supplied and its provenance is not recorded. Every profile hangs off it. - **Confirming DMA is actually used.** If transfers fall back to PIO the CPU cost rises far above the ~12-15% cycle-steal estimate and CPU becomes the binding constraint. This is the worst plausible outcome and the cheapest to check — do it first. **Do not try to get the bandwidth number out of MAME.** Its SCSI/SASI devices are functional models, not timing-accurate; a KB/s figure from MAME measures the emulator's scheduler. `docs/BENCHMARK.md` covers the three-tier approach (MAME validates the path, derivation bounds it, real hardware settles it). ## Display path — WORKING, verified end to end (session 3) The first real frame is on screen: `docs/images/x68k_first_frame_compare.png`. Full write-up in **FINDINGS 22**. Harness: `tools/bench/show_frame.lua` + `tools/bench/prep_frame.py`. Three facts the player MUST honour, none of which were guessable: | what | where | value | |---|---|---| | **Un-hide the graphics layer** | CRTC R20 `$E80028` | clear bit 11 ("G-VRAM set to buffer"); IPL leaves `0x0B16` | | Colour setup (256c) | CRTC R20 bits 9-8 | `0x0100` -> `R20 = 0x0116` | | **Monitor contrast** | `$E8E001` bits 3-0 | IPL leaves **14**; write **15** or everything renders 7% dark | Bit 11 is the one that cost the most time: GVRAM writes land and read back correctly while the layer is invisible, so the video controller looks guilty and is not. Contrast `0` blanks the screen — free fade-to-black for transitions. Palette format is now **confirmed from MAME source**, not assumed: `GGGGGRRRRRBBBBBI` (G 15:11, R 10:6, B 5:1, shared LSB I), expanded as `pal6bit((field<<1)|I)`. With contrast at 15 the render is **pixel-exact**. New ceiling: the 15-bit+I palette alone costs **38.88 dB** against the 24-bit palettised source — the same order as the `scsi` profile's own codec error (39.4 dB). `scsi` is close to display-transparent on real hardware. See FINDINGS 22.4 before considering raising quality further. Snapshot recipe that works (`-video none` CANNOT snapshot): ``` SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \ -sound none -nothrottle -plugins -autoboot_script