Session 2 reversed several of its own conclusions. The docs are append-only, so
a reader could land on a superseded section and act on it. This pass makes the
repo internally consistent.
Defects found and fixed in STATUS.md:
- claimed "Hybrid VQ with k=1024: no" as the answer to the linework question,
directly contradicting FINDINGS 14, which rejected k=1024. Both profiles are
k=256.
- malformed profile table (six column separators, five columns).
- next-steps list had two items numbered 3 and listed the full-disc survey
twice.
- the disk-benchmark section still read CRITICAL-PATH with "if SCSI sustains
>=800 KB/s, ship pixel-exact". That was written while the bandwidth figure
was misread as 4 MB/s. At 4 Mbps pixel-exact needs 92-97% of the pipe and is
not available, and the ring-buffer result means the design no longer hangs on
the benchmark at all. Rewritten with what it IS still worth doing: confirming
the 4 Mbps provenance, and confirming DMA is used rather than PIO.
FINDINGS now carries supersession blockquotes on 5, 8, 11, 17 and 18 pointing
at the sections that correct them. 18 is the dangerous one -- its peak-vs-
sustained test is reversed by 21 -- so it is marked DO NOT ACT ON THIS SECTION
while noting the per-frame data itself remains valid.
profile_gen.py had the same problem in code: it defaulted to the superseded
peak sizing and returned lam=25 where the docs say lam=10. The buffered test is
now the default and peak sizing is behind --size-for-peak as a bound only. A
tool that contradicts the findings is worse than no tool.
Also preserves the five measurement scripts that produced this session's
numbers as tools/analysis/05-09, following the session 1 precedent, and adds an
"explicitly abandoned -- do not re-propose" list to STATUS covering entropy
coding, k=1024 codebooks and flat 4x4 VQ.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
User clarified the bandwidth figure is 4 Mbps (488 KB/s), not 4 MB/s -- ~8x
tighter than the previous commit reasoned against. Two consequences, plus a
correction to session 1.
1. The scsi profile committed in f0f2f80 DOES NOT FIT. Its mean is a
comfortable 52% of the pipe but it PEAKS at 96.4% (470.8 KB/s on scene
00020), and a frame that arrives late is a dropped frame, not a slow one.
Peak/mean is 1.4-1.9x even on 1.2-1.7s clips. Sizing a real-time stream on
the mean was the error. Flagged in STATUS rather than silently retuned,
because the fix is rate control, not a lower lam.
This promotes ratectl.py -- written in session 2, never wired into
encode.py -- from a loose end to the highest-value work in the repo. It is
worth a full step on the quality ladder (lam=25 -> lam=10, +0.7/+1.2 dB)
because it allows sizing for the mean instead of the peak.
2. Pixel-exact is off the table at this bandwidth: lam=0 needs 92-97% of the
pipe. The previous commit's "if SCSI sustains >=800 KB/s, ship transparent"
conclusion only applies at roughly double the user's figure.
3. FINDINGS 5 said that because transfers are DMA, streaming "costs essentially
no CPU" and the 68000 is "nearly idle". That is wrong. The HD63450 steals
~8 clocks per 16-bit word: 10-20% of the machine at the rates the profiles
now use, on top of a 38% full-frame blit. Bandwidth and CPU are one budget.
Adds tools/encoder/profile_gen.py, which derives lam FROM a bandwidth figure
(accounting for audio, peak/mean and DMA steal) instead of reading it off the
knee of the RD curve, and docs/BENCHMARK.md covering how to actually measure
the subsystem -- including why MAME cannot answer the bandwidth question and
would be the same class of error as the FINDINGS 4 traps.
The 4 Mbps figure is user-supplied and its provenance is not recorded; every
profile now hangs off it.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
The profiles shipped in e4062ed were set far too low. 45 KB/s (sasi) and
75 KB/s (scsi) are 12% and 7% of the respective folklore bus figures. They had
been read off the knee of the rate-distortion curve and then presented as
though bandwidth-derived, which they were not.
Raised to sasi 110 KB/s (lam=60) and scsi 280 KB/s (lam=10) -- 35% and 28%
utilisation. scsi is now within 0.52 dB of the palette ceiling on scene 00020.
Checking the CPU side, which nobody had done for the decode path, produces a
second and more important result. Against the 833k cycle/frame budget at 12fps:
full-frame blit, every frame 319k 38% affordable
LZ4/LZSS decode ~30KB/frame 450k 54%
deflate decode ~30KB/frame 1800k 216% infeasible
So raising the VQ bitrate is nearly free -- RAW, the mode that dominates at
high rate, is the cheapest mode to blit -- but entropy coding is not viable at
all. That demotes the "247 KB/s lossless changed-spans+deflate" figure from
FINDINGS 8 to a compression upper bound rather than a shippable design, and
removes entropy coding from the roadmap. VQ is the right architecture precisely
because its decode is a table copy.
Also confirms the architecture unifies: the hybrid at lam=0 lands within 3% of
the purpose-built lossless coder, so there is no separate lossless path.
Consequence for planning: the blocked disk benchmark is now critical-path, not
optional. If SCSI sustains >=800 KB/s the correct scsi profile is lam=0 --
pixel-exact video at ~450 KB/s and 38% CPU. Whether this port ships transparent
or lossy on SCSI is waiting on one measurement.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Answers session 1's critical-path question. Flat 4x4 VQ at k=256 was prototyped
and REJECTED by eye: Dirk's face disintegrates and ink outlines break into
4-pixel stair-steps. The 256-colour palettised frame is excellent, so the
palette was never the problem -- block VQ was.
Replaced it with a Cinepak-style hybrid: each 4x4 block is SKIP, one 4x4
codeword, four 2x2 codewords, or RAW literal pixels, chosen per block by
rate-distortion. The RAW escape makes lam=0 pixel-exact (measured 0.00 dB loss),
so the quality knob spans lossless to heavily-compressed in one bitstream.
Per the user's decision, ships TWO quality profiles from that one codec, one
decoder and one bitstream -- only the rate knob differs:
sasi 45 KB/s lam=300 34.8 dB stock 10MHz ACE/EXPERT
scsi 75 KB/s lam=100 35.9 dB Super/XVI or CZ-6BS1
Three corrections to earlier numbers:
1. Session 1's "183 KB/s at 12fps" was a bad extrapolation. Halving the
framerate does not halve the bitrate -- decimation roughly doubles the
per-frame delta. Re-measured directly: 340 KB/s for session 1's own RLE,
247 KB/s for changed-spans+deflate. The lossless floor is 319 MB.
2. A FOURTH false-good result, same family as the three in FINDINGS 4:
k=1024 codebooks appeared to buy +2.4 dB free, because the rate model
charged 1 byte for a 10-bit index. Charging the true cost reverses the
verdict -- k=256 wins at every matched bitrate, and by 5 dB at the low end
where the SASI profile lives. k=256 ships.
3. Stream inventory: the ~3-5MB clips are 1.2-1.7s, not ~60s, and some 60s
streams are menus, not content. Any survey must classify before averaging.
Also cleared both candidate sources for the game-logic layer: the SNES project
is MIT and DirkSimple is zlib, so the arcade scene graph can be imported and
the two transcriptions diffed against each other.
Encoder is working end-to-end: extract.py -> vq/vq_hybrid/ratectl -> encode.py,
emitting a big-endian DLX1 container the 68000 can parse with plain moves.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
Verified GVRAM is one word-access per pixel in ALL color modes; chose 256-color
256x192 with movem.l bursts (page 1 sacrificed as double-buffer).
Measured 8 scenes from the Blu-ray source: blit costs under 8% of the 12fps
cycle budget, so I/O is the bottleneck, not CPU. Naive delta+RLE reaches only
3.2:1 (365 KB/s, 470MB) -> decision to use 4x4 vector quantization (~30 KB/s).
"Shot on twos" assumption failed: the transfer has zero duplicate frames, so
12fps requires explicit decimation.
Documents three false measurement results and their root causes (per-frame
Floyd-Steinberg dithering, temporal denoise, exact-match dedupe on noisy source).
MAME Lua injection harness works and is reusable for cycle-cost measurement;
the IOCS _B_READ disk benchmark is blocked returning -1.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6