The session-3 milestone was written in a way that reads as "the port renders", which it does not. The video hardware is genuinely emulated and the output is bit-exact, but GVRAM was filled by a MAME Lua script poking emulated memory, not by 68000 instructions. The distinction is load-bearing: Lua writes cost zero 68000 cycles, so nothing here tests whether the CPU can decode and blit inside 833,333 cycles. The 38% full-frame blit estimate that the entire budget rests on is still unvalidated. Only the "Not yet started" list carried this caveat, which was too buried for a claim this easy to over-read. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
304 lines
16 KiB
Markdown
304 lines
16 KiB
Markdown
# Status & next-session handoff — end of session 3 (2026-08-23)
|
|
|
|
## Decisions locked
|
|
|
|
| decision | value | why |
|
|
|---|---|---|
|
|
| Target CPU | 68000 @ 10MHz (stock) | hardest honest constraint |
|
|
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
|
|
| Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway |
|
|
| **Codec** | **hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion** | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 |
|
|
| **Quality modes** | **two: `sasi` and `scsi`** (USER DECISION, session 2) | one codec, one decoder, one bitstream; only `lam` differs |
|
|
| Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win |
|
|
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
|
|
| SNES project reuse | **MIT — cleared** | `data/events/` scene graph is reusable with attribution |
|
|
|
|
### The SASI/SCSI question is RESOLVED
|
|
Session 1 left "which machine do we target" open. The user's answer: **ship both**,
|
|
as two quality profiles. This is now implemented rather than hypothetical — the
|
|
bitrate ceiling is a build parameter in `tools/encoder/ratectl.py`:
|
|
|
|
| profile | target | lam | quality (00020 / 00146) | machine |
|
|
|---|---|---|---|---|
|
|
| `sasi` | 110 KB/s | 60 | 36.9 / 29.6 dB | stock 10MHz ACE/EXPERT |
|
|
| `scsi` | 280 KB/s | 10 | 39.4 / 32.3 dB | Super/XVI, or CZ-6BS1 board |
|
|
|
|
Sized against the user's working figure of **4 Mbps = 488 KB/s sustained**, on
|
|
SD-backed SCSI (BlueSCSI / SCSI2SD) — so that rate is a bus-limited **constant**,
|
|
not an average over seek latency.
|
|
|
|
**Both profiles fit with room.** Ring-buffer simulation on the real per-frame
|
|
sizes gives **zero required prefill** for every scene at both profiles: the fill
|
|
delivers 40.69 KB per frame time and only one measured frame (42.10 KB) exceeds
|
|
that, recovered by the next. A 256 KB buffer carries ~1 s of stall tolerance,
|
|
far more than an SD-backed seek needs. FINDINGS 21.
|
|
|
|
An earlier warning here said `scsi` did not fit because a frame peaked at 96.4%
|
|
of the pipe. That compared instantaneous demand to a sustained rate as if they
|
|
had to match frame-by-frame; with a buffer the test is cumulative, and it passes.
|
|
|
|
`scsi` is now within **0.5 dB of the palette ceiling** on 00020. These were
|
|
initially set at 45 / 75 KB/s, which was 12% / 7% bus utilisation — read off the
|
|
RD curve rather than derived from the hardware. See FINDINGS 17.
|
|
|
|
Codebooks are **k=256 with 1-byte indices** in both profiles. k=1024 was measured
|
|
and rejected — see FINDINGS 14, it was a false-good result from a rate model
|
|
that undercharged the index. Do not ship past `lam~800`; FINDINGS 15 has the cliff.
|
|
|
|
Because of the RAW escape mode, `lam=0` is **pixel-exact** against the palettised
|
|
frame (measured 0.00 dB loss). The profiles are two points on one continuous
|
|
rate-distortion curve, not two codecs.
|
|
|
|
---
|
|
|
|
## What session 3 settled
|
|
|
|
1. **The display path works and is verified end to end.** First real frame on an
|
|
emulated X68000 screen: `docs/images/x68k_first_frame_compare.png`. Full
|
|
write-up in **FINDINGS 22**. Everything before this session was Python-side
|
|
or a headless `-video none` run, which cannot snapshot at all.
|
|
2. **The render is pixel-exact, not merely close.** With monitor contrast at 15,
|
|
all 256 palette entries render exactly as `GGGGGRRRRRBBBBBI` + `pal6bit`
|
|
predicts. That exactness is the regression test — see
|
|
`tools/bench/verify_frame.py`, which exits non-zero if it ever drifts.
|
|
3. **Three hardware facts that were previously assumed are now confirmed from
|
|
MAME 0.277 source**, not folklore: the palette word format, the 1024-byte
|
|
GVRAM line stride, and the 256-colour page aliasing in `HARDWARE.md`. All
|
|
three were already written down correctly; they are now cited.
|
|
4. **A new quality ceiling was measured** — the 15-bit+I palette alone costs
|
|
38.88 dB, the same order as the `scsi` profile's own codec error. FINDINGS
|
|
22.4. This bounds how much further `scsi` is worth raising.
|
|
5. **Two shell traps that wedged session 2's background jobs** are documented in
|
|
the working-setup section below. They cost ~1.5 h of wall clock and a wedged
|
|
CPU core, and one of them was hit again this session.
|
|
|
|
## What session 2 settled
|
|
|
|
1. **The critical-path question is answered.** "Does VQ soften Bluth's linework
|
|
unacceptably?" — **flat 4x4 VQ: yes, badly. The hybrid (SKIP/V1/V4/RAW): no.**
|
|
Verified by eye, not just PSNR. See `docs/FINDINGS.md` 9-11 and the two
|
|
images in `docs/images/`. Both profiles use **k=256**; see item 2b.
|
|
2. **Session 1's 12fps bitrate was wrong** (183 KB/s claimed, 340 KB/s measured).
|
|
Halving the framerate does not halve the bitrate. FINDINGS 8.
|
|
2b. **A fourth false-good result was produced and caught this session** — k=1024
|
|
codebooks looked like a +2.4 dB free win because the rate model charged 1 byte
|
|
for a 10-bit index. FINDINGS 14. The k=256 configuration ships.
|
|
3. **The 256-colour palettised frame is the real quality ceiling** and it looks
|
|
excellent. Judge the codec against that, not against 1080p.
|
|
4. Encoder exists and produces a real bitstream: `tools/encoder/`.
|
|
|
|
---
|
|
|
|
## Encoder — working
|
|
|
|
```
|
|
python3 tools/encoder/extract.py 00020 /tmp/fr_00020 12 crop
|
|
python3 tools/encoder/encode.py /tmp/fr_00020 out.dlx --profile sasi --preview p.png
|
|
```
|
|
|
|
| file | role |
|
|
|---|---|
|
|
| `extract.py` | .m2ts -> 256x192 PNGs, 12fps, spatial-only denoise |
|
|
| `vq.py` | palette, blockify, hand-rolled k-means (no sklearn on this box), PSNR |
|
|
| `vq_hybrid.py` | the codec: 4 block modes + lagrangian mode decision |
|
|
| `ratectl.py` | SASI/SCSI profiles, leaky-bucket rate control |
|
|
| `encode.py` | CLI + `DLX1` container writer |
|
|
|
|
`DLX1` container layout is documented in the `encode.py` docstring. All
|
|
multi-byte fields are **big-endian** so the 68000 reads them with a plain `move`.
|
|
|
|
### Known encoder gaps
|
|
- **Rate control is written but not yet wired into `encode.py`** — the CLI uses a
|
|
fixed `lam` from the profile. `ratectl.encode_rate_controlled()` exists and
|
|
builds a lam-ladder per frame; it needs hooking up and validating.
|
|
- **Payload is deliberately NOT entropy-coded** — deflate decode does not fit in
|
|
the 68000's frame budget (FINDINGS 17.2). Do not "optimise" this later.
|
|
- Codebooks are per-scene and rebuilt from scratch; no inter-scene reuse.
|
|
- `_paint` is a Python per-block loop — fine for prototyping, slow for a full
|
|
disc encode. Vectorise before the 224-stream run.
|
|
|
|
---
|
|
|
|
## Working setup (unchanged from session 1, re-verified)
|
|
|
|
**MAME ROMs** — `~/mame/roms/x68000.zip`. Must pass **`-bios ipl10`**.
|
|
```
|
|
mame x68000 -bios ipl10 -video none -sound none -nothrottle -seconds_to_run 3
|
|
```
|
|
**Assembler** — `tools/vasm/vasmm68k_mot -Fbin -o out.bin in.s`
|
|
|
|
**Blu-ray** — `udisksctl loop-setup -r -f DRAGONS_LAIR.iso` -> `/media/reala-misaki/BDROM`
|
|
(still mounted as of end of session 2).
|
|
|
|
**MAME Lua harness** — `tools/bench/*.lua`, working. Three gotchas (retain the
|
|
notifier subscription in a global; the stack register is `SP` not `A7`;
|
|
`autoboot_script` fires at PC=0 before boot) are documented in FINDINGS.
|
|
|
|
**Two shell traps, both hit again this session:**
|
|
- piping MAME (or any long job) through `grep` block-buffers — write to a file.
|
|
- `pkill -f <pattern>` matches your own shell and kills it (exit 144).
|
|
Use `pkill -x` or kill by PID.
|
|
- **`until ! pgrep -f foo.py; do sleep; done` watcher loops never exit.** The
|
|
watching shell's own command line contains the string `foo.py`, so `pgrep -f`
|
|
matches the watcher itself and the loop spins forever. Session 2 left 11 of
|
|
these wedged for over an hour. Wait on the PID (`while kill -0 $PID`) or on a
|
|
sentinel file the job touches when it finishes -- never on a `-f` name match.
|
|
- **`timeout N mame ...` does not kill MAME.** MAME catches SIGTERM and, with an
|
|
autoboot script blocked waiting on a flag that never arrives, never reaches
|
|
its shutdown path. `timeout` without `-k` then waits forever while MAME burns
|
|
a full core at `-nothrottle`. Always `timeout -k 5 N`.
|
|
|
|
---
|
|
|
|
## Disk throughput benchmark — still blocked, no longer gating
|
|
|
|
`IOCS _B_READ` returns -1 uniformly. Full diagnosis and the four untested
|
|
hypotheses are in session 1's notes (git history of this file, commit 65112b9);
|
|
the ordered plan for retrying is in **`docs/BENCHMARK.md`**.
|
|
|
|
**Status changed twice this session — read this rather than the git history.**
|
|
It was briefly promoted to critical-path while the working bandwidth figure was
|
|
misread as 4 MB/s. With the correct figure (**4 Mbps = 488 KB/s**) and the
|
|
ring-buffer simulation showing **zero required prefill** for both profiles
|
|
(FINDINGS 21), the design no longer hangs on it. Pixel-exact on SCSI is **not**
|
|
available at 4 Mbps — it needs 92-97% of the pipe — so there is no longer a
|
|
"measure it and maybe ship transparent" decision waiting.
|
|
|
|
What the benchmark is still worth doing for:
|
|
- **Confirming the 4 Mbps figure.** It is user-supplied and its provenance is
|
|
not recorded. Every profile hangs off it.
|
|
- **Confirming DMA is actually used.** If transfers fall back to PIO the CPU
|
|
cost rises far above the ~12-15% cycle-steal estimate and CPU becomes the
|
|
binding constraint. This is the worst plausible outcome and the cheapest to
|
|
check — do it first.
|
|
|
|
**Do not try to get the bandwidth number out of MAME.** Its SCSI/SASI devices are
|
|
functional models, not timing-accurate; a KB/s figure from MAME measures the
|
|
emulator's scheduler. `docs/BENCHMARK.md` covers the three-tier approach
|
|
(MAME validates the path, derivation bounds it, real hardware settles it).
|
|
|
|
## Display path — VERIFIED (session 3). CPU path — still unproven.
|
|
|
|
The first real frame is on screen: `docs/images/x68k_first_frame_compare.png`.
|
|
|
|
**What this does and does not mean.** The video hardware is genuinely emulated
|
|
and the render is bit-exact. But GVRAM was filled by a MAME Lua script, not by
|
|
68000 code — no 68000 instruction has drawn a pixel yet. Lua writes cost zero
|
|
68000 cycles, so the 38% full-frame blit estimate underpinning the whole CPU
|
|
budget is still unvalidated. "Verified end to end" applies to the *display*
|
|
path only. See FINDINGS 22 scope note.
|
|
Full write-up in **FINDINGS 22**. Harness: `tools/bench/show_frame.lua` +
|
|
`tools/bench/prep_frame.py`.
|
|
|
|
Three facts the player MUST honour, none of which were guessable:
|
|
|
|
| what | where | value |
|
|
|---|---|---|
|
|
| **Un-hide the graphics layer** | CRTC R20 `$E80028` | clear bit 11 ("G-VRAM set to buffer"); IPL leaves `0x0B16` |
|
|
| Colour setup (256c) | CRTC R20 bits 9-8 | `0x0100` -> `R20 = 0x0116` |
|
|
| **Monitor contrast** | `$E8E001` bits 3-0 | IPL leaves **14**; write **15** or everything renders 7% dark |
|
|
|
|
Bit 11 is the one that cost the most time: GVRAM writes land and read back
|
|
correctly while the layer is invisible, so the video controller looks guilty and
|
|
is not. Contrast `0` blanks the screen — free fade-to-black for transitions.
|
|
|
|
Palette format is now **confirmed from MAME source**, not assumed:
|
|
`GGGGGRRRRRBBBBBI` (G 15:11, R 10:6, B 5:1, shared LSB I), expanded as
|
|
`pal6bit((field<<1)|I)`. With contrast at 15 the render is **pixel-exact**.
|
|
|
|
New ceiling: the 15-bit+I palette alone costs **38.88 dB** against the 24-bit
|
|
palettised source — the same order as the `scsi` profile's own codec error
|
|
(39.4 dB). `scsi` is close to display-transparent on real hardware. See
|
|
FINDINGS 22.4 before considering raising quality further.
|
|
|
|
Snapshot recipe that works (`-video none` CANNOT snapshot):
|
|
```
|
|
SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -video soft -window \
|
|
-sound none -nothrottle -plugins -autoboot_script <script>.lua \
|
|
-snapshot_directory ./snap -snapview native -seconds_to_run 6
|
|
```
|
|
`-snapview native` drops MAME's LED artwork and gives a clean 768x512 screen.
|
|
|
|
## Next steps, in priority order
|
|
|
|
1. **Full-disc survey.** Only 4 clips of 1.2-1.7 s out of 224 streams have been
|
|
measured, and 00146 already runs 23% hotter than 00020. A *sustained* action
|
|
sequence is the one thing that could still break the bitrate. Classify menu
|
|
vs content first (FINDINGS 13) or the averages are diluted by static menus.
|
|
**Vectorise `_paint` before this run** — it is a Python per-block loop.
|
|
2. **68000 decoder skeleton.** Parse `DLX1`, expand codebooks to word-per-pixel,
|
|
blit SKIP/V1/V4/RAW. Measure real cycles with the existing MAME Lua harness.
|
|
**Now unblocked** — the display path is verified (FINDINGS 22) and
|
|
`tools/bench/show_frame.lua` gives a known-good reference image to diff the
|
|
68000's output against. Validates the 38% full-frame blit estimate that the
|
|
whole CPU budget rests on. Still needs a real CRTC mode table for 256x256;
|
|
the harness deliberately borrows the IPL's timing and invents nothing.
|
|
2a. **CRTC mode table for 256x192-in-256x256.** Prerequisite for (2) and the
|
|
smallest well-defined unit of work available right now. Needs real R00-R08
|
|
timing values. **Do not write these from memory** — session 3 lost time to
|
|
exactly that failure mode on the video registers. Derive them from the CRTC
|
|
dividers in `x68k_crtc.cpp` (`m_reg[20] & 0x1f` selects the dot-clock
|
|
divisor; the IPL's `0x16` gives /2 off the 69MHz clock), or lift a known-good
|
|
set from a real X68000 title and verify by snapshot. The harness makes this
|
|
cheap to iterate: change values, snapshot, look.
|
|
|
|
3. **Wire rate control into `encode.py`.** No longer a blocker (FINDINGS 21), but
|
|
it is what gives a deterministic ceiling over content not yet measured, which
|
|
was the original reason for choosing VQ. Insurance, not a fix. Pairs with (1).
|
|
4. **Confirm DMA vs PIO in MAME** (see the benchmark section above) — cheap, and
|
|
the only thing that could still move CPU into the binding position.
|
|
5. **Resolve the framing question** (FINDINGS 12: crop vs squash vs wide).
|
|
Needs an eyeball against arcade reference, not a measurement.
|
|
6. **Import the scene graph.** SNES project `data/events/` (MIT, cleared),
|
|
cross-checked against DirkSimple (zlib) which transcribed the same data
|
|
independently — diff them to catch transcription errors before committing
|
|
any of it to 68000 tables.
|
|
7. **ADPCM audio.** MSM6258, 15.6kHz mono, 7.8 KB/s — already budgeted in
|
|
`ratectl.py`, not yet extracted or encoded.
|
|
|
|
### Explicitly abandoned — do not re-propose
|
|
- ~~Entropy-code the payload.~~ Deflate decode is ~216% of the frame budget on a
|
|
68000; LZ4 is ~54% with no room beside a 38% blit (FINDINGS 17.2). All bitrates
|
|
are raw payload. This also demotes the "247 KB/s lossless" figure in FINDINGS 8
|
|
to a compression upper bound, not a shippable design.
|
|
- ~~k=1024 codebooks.~~ False-good result from a rate model that charged 1 byte
|
|
for a 10-bit index (FINDINGS 14). k=256 wins at every matched bitrate.
|
|
- ~~Flat 4x4 VQ.~~ Rejected by eye (FINDINGS 9).
|
|
|
|
## Not yet started
|
|
- **Any 68000 player code.** `src/player/` is still empty. The display path is
|
|
proven, but proven *from Lua* — no 68000 instruction has yet drawn a pixel.
|
|
- **A real CRTC mode table.** The harness deliberately borrows the IPL's 768x512
|
|
text timing and invents no CRTC values, which is why the frame repeats at
|
|
x=512 (FINDINGS 22.5). A 256x256 mode needs real R00-R08 values, and those
|
|
must be derived or measured, NOT recalled from memory — see the note below.
|
|
- ADPCM audio extraction/encoding
|
|
- Disk image packaging
|
|
- Game logic (scene branching, input windows, death clips)
|
|
|
|
## Reproducing the display result
|
|
|
|
```
|
|
python3 tools/encoder/extract.py 00020 tmp/fr_00020 12 crop
|
|
python3 tools/bench/prep_frame.py tmp/fr_00020 tmp/frame.bin 0
|
|
mkdir -p tmp/snap_verify && cd tmp && SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 \
|
|
-video soft -window -sound none -nothrottle -plugins \
|
|
-autoboot_script ../tools/bench/show_frame.lua \
|
|
-snapshot_directory ./snap_verify -snapview native -seconds_to_run 6
|
|
cd .. && python3 tools/bench/verify_frame.py
|
|
```
|
|
Verified cold from the Blu-ray at end of session 3: exact match, 38.88 dB.
|
|
|
|
`tmp/` is gitignored scratch. The frames are NOT in the repo — regenerate them
|
|
with `extract.py`; the earlier ones lived in `/tmp` and do not survive a reboot.
|
|
|
|
## Reference material on this box (not in the repo)
|
|
|
|
- **MAME 0.277 source: `~/src/mame-mame0277/`** (tarball `~/src/mame0277.tar.gz`).
|
|
Downloaded this session to settle the graphics-layer question. The files that
|
|
matter are `src/mame/sharp/x68k_v.cpp`, `x68k_crtc.cpp`, `x68k_crtc.h`,
|
|
`x68k.cpp`. **Read these before theorising about X68000 video behaviour** —
|
|
six register-poking attempts failed against a gate that one grep found.
|
|
- Blu-ray mounted at `/media/reala-misaki/BDROM` via
|
|
`udisksctl loop-setup -r -f DRAGONS_LAIR.iso`.
|