Eleven watcher shells and one MAME instance were left running for over an hour. Both had the same shape: a wait that can never be satisfied. - `until ! pgrep -f foo.py` matches the watching shell's own command line, so the loop never terminates. Wait on a PID or a sentinel file instead. - `timeout N mame` sends a SIGTERM that MAME ignores when its autoboot script is blocked; without `-k` the process spins at 100% CPU forever. Also gitignore vasm's default `a.out` output. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
10 KiB
Status & next-session handoff — end of session 2 (2026-08-23)
Decisions locked
| decision | value | why |
|---|---|---|
| Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint |
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
| Double buffer | none — page 1 sacrificed | enables movem.l 24px bursts; delta coding needs a RAM reference frame anyway |
| Codec | hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 |
| Quality modes | two: sasi and scsi (USER DECISION, session 2) |
one codec, one decoder, one bitstream; only lam differs |
| Framerate | 12 fps, explicit decimation | source has zero duplicate frames; no free "twos" win |
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
| SNES project reuse | MIT — cleared | data/events/ scene graph is reusable with attribution |
The SASI/SCSI question is RESOLVED
Session 1 left "which machine do we target" open. The user's answer: ship both,
as two quality profiles. This is now implemented rather than hypothetical — the
bitrate ceiling is a build parameter in tools/encoder/ratectl.py:
| profile | target | lam | quality (00020 / 00146) | machine |
|---|---|---|---|---|
sasi |
110 KB/s | 60 | 36.9 / 29.6 dB | stock 10MHz ACE/EXPERT |
scsi |
280 KB/s | 10 | 39.4 / 32.3 dB | Super/XVI, or CZ-6BS1 board |
Sized against the user's working figure of 4 Mbps = 488 KB/s sustained, on SD-backed SCSI (BlueSCSI / SCSI2SD) — so that rate is a bus-limited constant, not an average over seek latency.
Both profiles fit with room. Ring-buffer simulation on the real per-frame sizes gives zero required prefill for every scene at both profiles: the fill delivers 40.69 KB per frame time and only one measured frame (42.10 KB) exceeds that, recovered by the next. A 256 KB buffer carries ~1 s of stall tolerance, far more than an SD-backed seek needs. FINDINGS 21.
An earlier warning here said scsi did not fit because a frame peaked at 96.4%
of the pipe. That compared instantaneous demand to a sustained rate as if they
had to match frame-by-frame; with a buffer the test is cumulative, and it passes.
scsi is now within 0.5 dB of the palette ceiling on 00020. These were
initially set at 45 / 75 KB/s, which was 12% / 7% bus utilisation — read off the
RD curve rather than derived from the hardware. See FINDINGS 17.
Codebooks are k=256 with 1-byte indices in both profiles. k=1024 was measured
and rejected — see FINDINGS 14, it was a false-good result from a rate model
that undercharged the index. Do not ship past lam~800; FINDINGS 15 has the cliff.
Because of the RAW escape mode, lam=0 is pixel-exact against the palettised
frame (measured 0.00 dB loss). The profiles are two points on one continuous
rate-distortion curve, not two codecs.
What session 2 settled
- The critical-path question is answered. "Does VQ soften Bluth's linework
unacceptably?" — flat 4x4 VQ: yes, badly. The hybrid (SKIP/V1/V4/RAW): no.
Verified by eye, not just PSNR. See
docs/FINDINGS.md9-11 and the two images indocs/images/. Both profiles use k=256; see item 2b. - Session 1's 12fps bitrate was wrong (183 KB/s claimed, 340 KB/s measured). Halving the framerate does not halve the bitrate. FINDINGS 8. 2b. A fourth false-good result was produced and caught this session — k=1024 codebooks looked like a +2.4 dB free win because the rate model charged 1 byte for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
- The 256-colour palettised frame is the real quality ceiling and it looks excellent. Judge the codec against that, not against 1080p.
- Encoder exists and produces a real bitstream:
tools/encoder/.
Encoder — working
python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
python3 tools/encoder/encode.py /tmp/fr_00020 out.dlx --profile sasi --preview p.png
| file | role |
|---|---|
extract.py |
.m2ts -> 256x192 PNGs, 12fps, spatial-only denoise |
vq.py |
palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR |
vq_hybrid.py |
the codec: 4 block modes + lagrangian mode decision |
ratectl.py |
SASI/SCSI profiles, leaky-bucket rate control |
encode.py |
CLI + DLX1 container writer |
DLX1 container layout is documented in the encode.py docstring. All
multi-byte fields are big-endian so the 68000 reads them with a plain move.
Known encoder gaps
- Rate control is written but not yet wired into
encode.py— the CLI uses a fixedlamfrom the profile.ratectl.encode_rate_controlled()exists and builds a lam-ladder per frame; it needs hooking up and validating. - Payload is deliberately NOT entropy-coded — deflate decode does not fit in the 68000's frame budget (FINDINGS 17.2). Do not "optimise" this later.
- Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
_paintis a Python per-block loop — fine for prototyping, slow for a full disc encode. Vectorise before the 224-stream run.
Working setup (unchanged from session 1, re-verified)
MAME ROMs — ~/mame/roms/x68000.zip. Must pass -bios ipl10.
mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3
Assembler — tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s
Blu-ray — udisksctl loop-setup -r -f DRAGONS_LAIR.iso -> /media/reala-misaki/BDROM
(still mounted as of end of session 2).
MAME Lua harness — tools/bench/*.lua, working. Three gotchas (retain the
notifier subscription in a global; the stack register is SP not A7;
autoboot_script fires at PC=0 before boot) are documented in FINDINGS.
Two shell traps, both hit again this session:
- piping MAME (or any long job) through
grepblock-buffers — write to a file. pkill -f <pattern>matches your own shell and kills it (exit 144). Usepkill -xor kill by PID.until ! pgrep -f foo.py; do sleep; donewatcher loops never exit. The watching shell's own command line contains the stringfoo.py, sopgrep -fmatches the watcher itself and the loop spins forever. Session 2 left 11 of these wedged for over an hour. Wait on the PID (while kill -0 $PID) or on a sentinel file the job touches when it finishes -- never on a-fname match.timeout N mame ...does not kill MAME. MAME catches SIGTERM and, with an autoboot script blocked waiting on a flag that never arrives, never reaches its shutdown path.timeoutwithout-kthen waits forever while MAME burns a full core at-nothrottle. Alwaystimeout -k 5 N.
Disk throughput benchmark — still blocked, no longer gating
IOCS _B_READ returns -1 uniformly. Full diagnosis and the four untested
hypotheses are in session 1's notes (git history of this file, commit 65112b9);
the ordered plan for retrying is in docs/BENCHMARK.md.
Status changed twice this session — read this rather than the git history. It was briefly promoted to critical-path while the working bandwidth figure was misread as 4 MB/s. With the correct figure (4 Mbps = 488 KB/s) and the ring-buffer simulation showing zero required prefill for both profiles (FINDINGS 21), the design no longer hangs on it. Pixel-exact on SCSI is not available at 4 Mbps — it needs 92-97% of the pipe — so there is no longer a "measure it and maybe ship transparent" decision waiting.
What the benchmark is still worth doing for:
- Confirming the 4 Mbps figure. It is user-supplied and its provenance is not recorded. Every profile hangs off it.
- Confirming DMA is actually used. If transfers fall back to PIO the CPU cost rises far above the ~12-15% cycle-steal estimate and CPU becomes the binding constraint. This is the worst plausible outcome and the cheapest to check — do it first.
Do not try to get the bandwidth number out of MAME. Its SCSI/SASI devices are
functional models, not timing-accurate; a KB/s figure from MAME measures the
emulator's scheduler. docs/BENCHMARK.md covers the three-tier approach
(MAME validates the path, derivation bounds it, real hardware settles it).
Next steps, in priority order
- Full-disc survey. Only 4 clips of 1.2-1.7 s out of 224 streams have been
measured, and 00146 already runs 23% hotter than 00020. A sustained action
sequence is the one thing that could still break the bitrate. Classify menu
vs content first (FINDINGS 13) or the averages are diluted by static menus.
Vectorise
_paintbefore this run — it is a Python per-block loop. - 68000 decoder skeleton. Parse
DLX1, expand codebooks to word-per-pixel, blit SKIP/V1/V4/RAW. Measure real cycles with the existing MAME Lua harness — the first time that harness gets used for its actual purpose. Validates the 38% full-frame blit estimate that the whole CPU budget rests on. - Wire rate control into
encode.py. No longer a blocker (FINDINGS 21), but it is what gives a deterministic ceiling over content not yet measured, which was the original reason for choosing VQ. Insurance, not a fix. Pairs with (1). - Confirm DMA vs PIO in MAME (see the benchmark section above) — cheap, and the only thing that could still move CPU into the binding position.
- Resolve the framing question (FINDINGS 12: crop vs squash vs wide). Needs an eyeball against arcade reference, not a measurement.
- Import the scene graph. SNES project
data/events/(MIT, cleared), cross-checked against DirkSimple (zlib) which transcribed the same data independently — diff them to catch transcription errors before committing any of it to 68000 tables. - ADPCM audio. MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in
ratectl.py, not yet extracted or encoded.
Explicitly abandoned — do not re-propose
Entropy-code the payload.Deflate decode is ~216% of the frame budget on a 68000; LZ4 is ~54% with no room beside a 38% blit (FINDINGS 17.2). All bitrates are raw payload. This also demotes the "247 KB/s lossless" figure in FINDINGS 8 to a compression upper bound, not a shippable design.k=1024 codebooks.False-good result from a rate model that charged 1 byte for a 10-bit index (FINDINGS 14). k=256 wins at every matched bitrate.Flat 4x4 VQ.Rejected by eye (FINDINGS 9).
Not yet started
- Any 68000 player code
- ADPCM audio extraction/encoding
- Disk image packaging
- Game logic (scene branching, input windows, death clips)