The session-3 milestone was written in a way that reads as "the port renders", which it does not. The video hardware is genuinely emulated and the output is bit-exact, but GVRAM was filled by a MAME Lua script poking emulated memory, not by 68000 instructions. The distinction is load-bearing: Lua writes cost zero 68000 cycles, so nothing here tests whether the CPU can decode and blit inside 833,333 cycles. The 38% full-frame blit estimate that the entire budget rests on is still unvalidated. Only the "Not yet started" list carried this caveat, which was too buried for a claim this easy to over-read. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
16 KiB
Status & next-session handoff — end of session 3 (2026-08-23)
Decisions locked
| decision | value | why |
|---|---|---|
| Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint |
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
| Double buffer | none — page 1 sacrificed | enables movem.l 24px bursts; delta coding needs a RAM reference frame anyway |
| Codec | hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 |
| Quality modes | two: sasi and scsi (USER DECISION, session 2) |
one codec, one decoder, one bitstream; only lam differs |
| Framerate | 12 fps, explicit decimation | source has zero duplicate frames; no free "twos" win |
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
| SNES project reuse | MIT — cleared | data/events/ scene graph is reusable with attribution |
The SASI/SCSI question is RESOLVED
Session 1 left "which machine do we target" open. The user's answer: ship both,
as two quality profiles. This is now implemented rather than hypothetical — the
bitrate ceiling is a build parameter in tools/encoder/ratectl.py:
| profile | target | lam | quality (00020 / 00146) | machine |
|---|---|---|---|---|
sasi |
110 KB/s | 60 | 36.9 / 29.6 dB | stock 10MHz ACE/EXPERT |
scsi |
280 KB/s | 10 | 39.4 / 32.3 dB | Super/XVI, or CZ-6BS1 board |
Sized against the user's working figure of 4 Mbps = 488 KB/s sustained, on SD-backed SCSI (BlueSCSI / SCSI2SD) — so that rate is a bus-limited constant, not an average over seek latency.
Both profiles fit with room. Ring-buffer simulation on the real per-frame sizes gives zero required prefill for every scene at both profiles: the fill delivers 40.69 KB per frame time and only one measured frame (42.10 KB) exceeds that, recovered by the next. A 256 KB buffer carries ~1 s of stall tolerance, far more than an SD-backed seek needs. FINDINGS 21.
An earlier warning here said scsi did not fit because a frame peaked at 96.4%
of the pipe. That compared instantaneous demand to a sustained rate as if they
had to match frame-by-frame; with a buffer the test is cumulative, and it passes.
scsi is now within 0.5 dB of the palette ceiling on 00020. These were
initially set at 45 / 75 KB/s, which was 12% / 7% bus utilisation — read off the
RD curve rather than derived from the hardware. See FINDINGS 17.
Codebooks are k=256 with 1-byte indices in both profiles. k=1024 was measured
and rejected — see FINDINGS 14, it was a false-good result from a rate model
that undercharged the index. Do not ship past lam~800; FINDINGS 15 has the cliff.
Because of the RAW escape mode, lam=0 is pixel-exact against the palettised
frame (measured 0.00 dB loss). The profiles are two points on one continuous
rate-distortion curve, not two codecs.
What session 3 settled
- The display path works and is verified end to end. First real frame on an
emulated X68000 screen:
docs/images/x68k_first_frame_compare.png. Full write-up in FINDINGS 22. Everything before this session was Python-side or a headless-video nonerun, which cannot snapshot at all. - The render is pixel-exact, not merely close. With monitor contrast at 15,
all 256 palette entries render exactly as
GGGGGRRRRRBBBBBI+pal6bitpredicts. That exactness is the regression test — seetools/bench/verify_frame.py, which exits non-zero if it ever drifts. - Three hardware facts that were previously assumed are now confirmed from
MAME 0.277 source, not folklore: the palette word format, the 1024-byte
GVRAM line stride, and the 256-colour page aliasing in
HARDWARE.md. All three were already written down correctly; they are now cited. - A new quality ceiling was measured — the 15-bit+I palette alone costs
38.88 dB, the same order as the
scsiprofile's own codec error. FINDINGS 22.4. This bounds how much furtherscsiis worth raising. - Two shell traps that wedged session 2's background jobs are documented in the working-setup section below. They cost ~1.5 h of wall clock and a wedged CPU core, and one of them was hit again this session.
What session 2 settled
- The critical-path question is answered. "Does VQ soften Bluth's linework
unacceptably?" — flat 4x4 VQ: yes, badly. The hybrid (SKIP/V1/V4/RAW): no.
Verified by eye, not just PSNR. See
docs/FINDINGS.md9-11 and the two images indocs/images/. Both profiles use k=256; see item 2b. - Session 1's 12fps bitrate was wrong (183 KB/s claimed, 340 KB/s measured). Halving the framerate does not halve the bitrate. FINDINGS 8. 2b. A fourth false-good result was produced and caught this session — k=1024 codebooks looked like a +2.4 dB free win because the rate model charged 1 byte for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
- The 256-colour palettised frame is the real quality ceiling and it looks excellent. Judge the codec against that, not against 1080p.
- Encoder exists and produces a real bitstream:
tools/encoder/.
Encoder — working
python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
python3 tools/encoder/encode.py /tmp/fr_00020 out.dlx --profile sasi --preview p.png
| file | role |
|---|---|
extract.py |
.m2ts -> 256x192 PNGs, 12fps, spatial-only denoise |
vq.py |
palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR |
vq_hybrid.py |
the codec: 4 block modes + lagrangian mode decision |
ratectl.py |
SASI/SCSI profiles, leaky-bucket rate control |
encode.py |
CLI + DLX1 container writer |
DLX1 container layout is documented in the encode.py docstring. All
multi-byte fields are big-endian so the 68000 reads them with a plain move.
Known encoder gaps
- Rate control is written but not yet wired into
encode.py— the CLI uses a fixedlamfrom the profile.ratectl.encode_rate_controlled()exists and builds a lam-ladder per frame; it needs hooking up and validating. - Payload is deliberately NOT entropy-coded — deflate decode does not fit in the 68000's frame budget (FINDINGS 17.2). Do not "optimise" this later.
- Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
_paintis a Python per-block loop — fine for prototyping, slow for a full disc encode. Vectorise before the 224-stream run.
Working setup (unchanged from session 1, re-verified)
MAME ROMs — ~/mame/roms/x68000.zip. Must pass -bios ipl10.
mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3
Assembler — tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s
Blu-ray — udisksctl loop-setup -r -f DRAGONS_LAIR.iso -> /media/reala-misaki/BDROM
(still mounted as of end of session 2).
MAME Lua harness — tools/bench/*.lua, working. Three gotchas (retain the
notifier subscription in a global; the stack register is SP not A7;
autoboot_script fires at PC=0 before boot) are documented in FINDINGS.
Two shell traps, both hit again this session:
- piping MAME (or any long job) through
grepblock-buffers — write to a file. pkill -f <pattern>matches your own shell and kills it (exit 144). Usepkill -xor kill by PID.until ! pgrep -f foo.py; do sleep; donewatcher loops never exit. The watching shell's own command line contains the stringfoo.py, sopgrep -fmatches the watcher itself and the loop spins forever. Session 2 left 11 of these wedged for over an hour. Wait on the PID (while kill -0 $PID) or on a sentinel file the job touches when it finishes -- never on a-fname match.timeout N mame ...does not kill MAME. MAME catches SIGTERM and, with an autoboot script blocked waiting on a flag that never arrives, never reaches its shutdown path.timeoutwithout-kthen waits forever while MAME burns a full core at-nothrottle. Alwaystimeout -k 5 N.
Disk throughput benchmark — still blocked, no longer gating
IOCS _B_READ returns -1 uniformly. Full diagnosis and the four untested
hypotheses are in session 1's notes (git history of this file, commit 65112b9);
the ordered plan for retrying is in docs/BENCHMARK.md.
Status changed twice this session — read this rather than the git history. It was briefly promoted to critical-path while the working bandwidth figure was misread as 4 MB/s. With the correct figure (4 Mbps = 488 KB/s) and the ring-buffer simulation showing zero required prefill for both profiles (FINDINGS 21), the design no longer hangs on it. Pixel-exact on SCSI is not available at 4 Mbps — it needs 92-97% of the pipe — so there is no longer a "measure it and maybe ship transparent" decision waiting.
What the benchmark is still worth doing for:
- Confirming the 4 Mbps figure. It is user-supplied and its provenance is not recorded. Every profile hangs off it.
- Confirming DMA is actually used. If transfers fall back to PIO the CPU cost rises far above the ~12-15% cycle-steal estimate and CPU becomes the binding constraint. This is the worst plausible outcome and the cheapest to check — do it first.
Do not try to get the bandwidth number out of MAME. Its SCSI/SASI devices are
functional models, not timing-accurate; a KB/s figure from MAME measures the
emulator's scheduler. docs/BENCHMARK.md covers the three-tier approach
(MAME validates the path, derivation bounds it, real hardware settles it).
Display path — VERIFIED (session 3). CPU path — still unproven.
The first real frame is on screen: docs/images/x68k_first_frame_compare.png.
What this does and does not mean. The video hardware is genuinely emulated
and the render is bit-exact. But GVRAM was filled by a MAME Lua script, not by
68000 code — no 68000 instruction has drawn a pixel yet. Lua writes cost zero
68000 cycles, so the 38% full-frame blit estimate underpinning the whole CPU
budget is still unvalidated. "Verified end to end" applies to the display
path only. See FINDINGS 22 scope note.
Full write-up in FINDINGS 22. Harness: tools/bench/show_frame.lua +
tools/bench/prep_frame.py.
Three facts the player MUST honour, none of which were guessable:
| what | where | value |
|---|---|---|
| Un-hide the graphics layer | CRTC R20 $E80028 |
clear bit 11 ("G-VRAM set to buffer"); IPL leaves 0x0B16 |
| Colour setup (256c) | CRTC R20 bits 9-8 | 0x0100 -> R20 = 0x0116 |
| Monitor contrast | $E8E001 bits 3-0 |
IPL leaves 14; write 15 or everything renders 7% dark |
Bit 11 is the one that cost the most time: GVRAM writes land and read back
correctly while the layer is invisible, so the video controller looks guilty and
is not. Contrast 0 blanks the screen — free fade-to-black for transitions.
Palette format is now confirmed from MAME source, not assumed:
GGGGGRRRRRBBBBBI (G 15:11, R 10:6, B 5:1, shared LSB I), expanded as
pal6bit((field<<1)|I). With contrast at 15 the render is pixel-exact.
New ceiling: the 15-bit+I palette alone costs 38.88 dB against the 24-bit
palettised source — the same order as the scsi profile's own codec error
(39.4 dB). scsi is close to display-transparent on real hardware. See
FINDINGS 22.4 before considering raising quality further.
Snapshot recipe that works (-video none CANNOT snapshot):
SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
-sound none -nothrottle -plugins -autoboot_script <script>.lua \
-snapshot_directory ./snap -snapview native -seconds_to_run 6
-snapview native drops MAME's LED artwork and gives a clean 768x512 screen.
Next steps, in priority order
-
Full-disc survey. Only 4 clips of 1.2-1.7 s out of 224 streams have been measured, and 00146 already runs 23% hotter than 00020. A sustained action sequence is the one thing that could still break the bitrate. Classify menu vs content first (FINDINGS 13) or the averages are diluted by static menus. Vectorise
_paintbefore this run — it is a Python per-block loop. -
68000 decoder skeleton. Parse
DLX1, expand codebooks to word-per-pixel, blit SKIP/V1/V4/RAW. Measure real cycles with the existing MAME Lua harness. Now unblocked — the display path is verified (FINDINGS 22) andtools/bench/show_frame.luagives a known-good reference image to diff the 68000's output against. Validates the 38% full-frame blit estimate that the whole CPU budget rests on. Still needs a real CRTC mode table for 256x256; the harness deliberately borrows the IPL's timing and invents nothing. 2a. CRTC mode table for 256x192-in-256x256. Prerequisite for (2) and the smallest well-defined unit of work available right now. Needs real R00-R08 timing values. Do not write these from memory — session 3 lost time to exactly that failure mode on the video registers. Derive them from the CRTC dividers inx68k_crtc.cpp(m_reg[20] & 0x1fselects the dot-clock divisor; the IPL's0x16gives /2 off the 69MHz clock), or lift a known-good set from a real X68000 title and verify by snapshot. The harness makes this cheap to iterate: change values, snapshot, look. -
Wire rate control into
encode.py. No longer a blocker (FINDINGS 21), but it is what gives a deterministic ceiling over content not yet measured, which was the original reason for choosing VQ. Insurance, not a fix. Pairs with (1). -
Confirm DMA vs PIO in MAME (see the benchmark section above) — cheap, and the only thing that could still move CPU into the binding position.
-
Resolve the framing question (FINDINGS 12: crop vs squash vs wide). Needs an eyeball against arcade reference, not a measurement.
-
Import the scene graph. SNES project
data/events/(MIT, cleared), cross-checked against DirkSimple (zlib) which transcribed the same data independently — diff them to catch transcription errors before committing any of it to 68000 tables. -
ADPCM audio. MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in
ratectl.py, not yet extracted or encoded.
Explicitly abandoned — do not re-propose
Entropy-code the payload.Deflate decode is ~216% of the frame budget on a 68000; LZ4 is ~54% with no room beside a 38% blit (FINDINGS 17.2). All bitrates are raw payload. This also demotes the "247 KB/s lossless" figure in FINDINGS 8 to a compression upper bound, not a shippable design.k=1024 codebooks.False-good result from a rate model that charged 1 byte for a 10-bit index (FINDINGS 14). k=256 wins at every matched bitrate.Flat 4x4 VQ.Rejected by eye (FINDINGS 9).
Not yet started
- Any 68000 player code.
src/player/is still empty. The display path is proven, but proven from Lua — no 68000 instruction has yet drawn a pixel. - A real CRTC mode table. The harness deliberately borrows the IPL's 768x512 text timing and invents no CRTC values, which is why the frame repeats at x=512 (FINDINGS 22.5). A 256x256 mode needs real R00-R08 values, and those must be derived or measured, NOT recalled from memory — see the note below.
- ADPCM audio extraction/encoding
- Disk image packaging
- Game logic (scene branching, input windows, death clips)
Reproducing the display result
python3 tools/encoder/extract.py 00020 tmp/fr_00020 12 crop
python3 tools/bench/prep_frame.py tmp/fr_00020 tmp/frame.bin 0
mkdir -p tmp/snap_verify && cd tmp && SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 \
-video soft -window -sound none -nothrottle -plugins \
-autoboot_script ../tools/bench/show_frame.lua \
-snapshot_directory ./snap_verify -snapview native -seconds_to_run 6
cd .. && python3 tools/bench/verify_frame.py
Verified cold from the Blu-ray at end of session 3: exact match, 38.88 dB.
tmp/ is gitignored scratch. The frames are NOT in the repo — regenerate them
with extract.py; the earlier ones lived in /tmp and do not survive a reboot.
Reference material on this box (not in the repo)
- MAME 0.277 source:
~/src/mame-mame0277/(tarball~/src/mame0277.tar.gz). Downloaded this session to settle the graphics-layer question. The files that matter aresrc/mame/sharp/x68k_v.cpp,x68k_crtc.cpp,x68k_crtc.h,x68k.cpp. Read these before theorising about X68000 video behaviour — six register-poking attempts failed against a gate that one grep found. - Blu-ray mounted at
/media/reala-misaki/BDROMviaudisksctl loop-setup -r -f DRAGONS_LAIR.iso.