Drop SASI on capacity, then find the budget never had the disk in it

USER DECISION: drop the `sasi` profile. Not on bandwidth -- on capacity. A SASI
volume is 40 MB, and the 22.8 min of unique scene footage on the source Blu-ray
(streams 00000-00201, measured, not recalled) is 146 MiB at the LOWEST rate this
codec makes -- more than the machine's whole 4-unit SASI space. `scsi` is the
only profile now. FINDINGS 32.

Then the user asked whether we were drawing the wrong conclusions about PIO vs
DMA, and we were, more broadly than the question implied. Every CPU figure in
FINDINGS 24-34 is scored against the full 833,333 cycles/frame with nothing
subtracted for moving the bitstream off disk. Debiting the HD63450 cycle-steal
at the long-standing 8 clk/word ESTIMATE, "1 frame of 120 misses" becomes 84 of
120, median 112.4%. PIO at the span rate is 99.8% of the machine. Spans buy
cycles by spending bandwidth and the bandwidth returns as steal, so 31.6's "fits
completely" becomes a worst frame of 114.3%. 10 fps absorbs it: median 93.7%,
1/120. FINDINGS 35. `11_cpu_budget.py` takes --io dma|pio|none, defaults to dma,
and warns if asked for none.

Also landed:
- item 1 done: the cost model checked against the 68000 on a cost-aware
  container, -3.07% to +0.01%, whole-window mean -1.22%. FINDINGS 34.
- item 4 done: the container carries its own 4-byte record alignment (DLX2).
  94/120 record starts were on odd addresses -- an address error, not a slow
  read -- now 0/120 for 16 B/s. Re-encoding reproduces 31.1 exactly. FINDINGS 33.
- a `scsi` window does not fit the 2 MB machine the rig emulates (2.84 MB of
  stream past a 0x200000 ceiling). The gate now verifies 80 of 120 frames and
  SAYS so, and fails loudly when the pass does not complete, instead of
  reporting a phantom 49,005-pixel diff. FINDINGS 36.

Three near-misses this session had one shape: an unobservable run nearly
produced a false finding. stdbuf -oL on any MAME job that prints progress -- a
file is block-buffered too, and a run that is merely finishing looks exactly
like one that is wedged.

check.sh ALL GREEN.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
prosolis
2026-08-23 17:09:47 -07:00
parent 06b98d4b47
commit 7d365b3ff5
12 changed files with 664 additions and 129 deletions
+19 -10
View File
@@ -1,6 +1,6 @@
# Dragon's Lair — Sharp X68000 port # Dragon's Lair — Sharp X68000 port
Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SASI/SCSI). Porting Dragon's Lair to a stock X68000 (68000 @ 10MHz, 2MB, SCSI).
This is fundamentally a **video codec problem**, not a game-logic problem: the This is fundamentally a **video codec problem**, not a game-logic problem: the
game logic is a scene table with branching input windows; the difficulty is game logic is a scene table with branching input windows; the difficulty is
@@ -55,10 +55,13 @@ tools/bench/ MAME Lua injection harness + 68000 benchmark sources.
`prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and `prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and
verify `src/player/decode.s`; the verify pass is in check.sh. verify `src/player/decode.s`; the verify pass is in check.sh.
tools/vasm/ vasm m68k assembler (built from source) tools/vasm/ vasm m68k assembler (built from source)
tools/encoder/ hybrid VQ encoder + DLX1 container writer (working). tools/encoder/ hybrid VQ encoder + DLX2 container writer (working).
DLX2 4-byte-aligns every frame record: an odd `move.l` is an
ADDRESS ERROR on a 68000, not a slow read (FINDINGS 28.3).
dlx.py is the reference DECODER -- ground truth for the 68000. dlx.py is the reference DECODER -- ground truth for the 68000.
src/player/ decode.s: the 68000 DLX1 decoder. Pixel-exact, and 31% of src/player/ decode.s: the 68000 DLX decoder. Pixel-exact; 1 frame of 120
frames over the 12fps CPU budget. See FINDINGS 28. over the 12fps CPU budget once the mode decision prices
cycles. See FINDINGS 28 and 31.
assets/ extracted frames/audio (gitignored) assets/ extracted frames/audio (gitignored)
``` ```
@@ -66,18 +69,24 @@ assets/ extracted frames/audio (gitignored)
``` ```
python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop python3 tools/encoder/extract.py 00020 /tmp/fr 12 crop
python3 tools/encoder/encode.py /tmp/fr out.dlx --profile sasi --preview p.png python3 tools/encoder/encode.py /tmp/fr out.dlx --profile scsi --preview p.png
``` ```
Two quality profiles ship from one codec and one decoder — `sasi` (110 KB/s) and **One profile: `scsi`, 280 KB/s.** The 110 KB/s `sasi` profile was dropped in
`scsi` (280 KB/s) are two points on the same rate-distortion curve. Both are session 9 on capacity, not bandwidth — a SASI volume is limited to 40 MB, and
**ceilings**: lam is bisected per frame under a leaky bucket, so the profile's the game's 22.8 minutes of footage is 146 MiB even at that rate (FINDINGS 32).
`lam` is a quality floor rather than a setting (`--fixed-lam` opts out). The rate point may return under another name once the delivery medium is
settled, because a 1x CD-ROM sustains ~150 KB/s and CD-ROM is the only period
medium with the capacity.
The profile bitrate is a **ceiling**: lam is bisected per frame under a leaky
bucket, so the profile's `lam` is a quality floor rather than a setting
(`--fixed-lam` opts out).
There are **two** ceilings, on two different axes. The second is the 68000's There are **two** ceilings, on two different axes. The second is the 68000's
decode budget: `mu` is bisected per frame against 833,333 cycles so the frame decode budget: `mu` is bisected per frame against 833,333 cycles so the frame
also *decodes* in time, which takes the worst sustained window from 37 frames also *decodes* in time, which takes the worst sustained window from 37 frames
over budget to 1 for 0.26 dB (FINDINGS 31). It is on by default; `--no-cpu-fit` over budget to 1 for 0.62 dB at `scsi` (FINDINGS 31). It is on by default; `--no-cpu-fit`
restores session 7 behaviour. Unlike bytes, cycles have no bucket — there is no restores session 7 behaviour. Unlike bytes, cycles have no bucket — there is no
double buffer to decode ahead into, so it is a hard per-frame ceiling. The codec is double buffer to decode ahead into, so it is a hard per-frame ceiling. The codec is
a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four a Cinepak-style hybrid: each 4x4 block is coded as SKIP, one 4x4 codeword, four
+346
View File
@@ -1783,3 +1783,349 @@ The intra frame lands at 91.5% — spans are what make a full redraw fit, which
30.6's arithmetic arriving in a real container. That row is still a **model** of 30.6's arithmetic arriving in a real container. That row is still a **model** of
a bitstream nothing implements; the two levers have never run on the 68000 a bitstream nothing implements; the two levers have never run on the 68000
together, and the ring-buffer question of 30.7 gets sharper at 449 KB/s. together, and the ring-buffer question of 30.7 gets sharper at 449 KB/s.
## 32. SASI is dropped, and the reason is capacity, not bandwidth (session 9)
**USER DECISION**: drop the `sasi` profile. A SASI volume on this machine is
limited to 40 MB, and the game does not fit in one.
That ends the two-quality-mode decision of session 2. `scsi` is now the only
profile, and `encode.py --profile` has one choice. The retired 110 KB/s rate
point is not deleted from the record, for the reason in 32.3.
### 32.1 How much video there actually is
Measured off the source Blu-ray rather than recalled: the unique scene footage
is streams `00000`-`00201`, **1366.6 s = 22.8 min**. The longer streams
(`00215` 1376 s, `00216` 1152 s, `00223` 566 s) are compilations of the same
material and are not additional content — 00223 is the window every codec
measurement in this project has been taken on. Total across all 224 streams is
88.3 min, which is the figure to *not* quote.
22.8 min agrees with the ~22 min of laserdisc footage the arcade original is
usually credited with, which is the cross-check that the compilations really
are duplicates.
At the rates this codec has actually produced, including the 7.8 KB/s audio
allowance:
| stream | rate | whole game |
|---|---:|---:|
| retired 110 KB/s profile | 109.4 KB/s | **146.0 MiB** |
| `scsi`, measured (FINDINGS 31) | 278.6 KB/s | **371.8 MiB** |
| `scsi` + spans (MODEL, 31.6) | 449.3 KB/s | **599.6 MiB** |
### 32.2 Where the 40 MB actually comes from
It is not a bus-addressing limit. MAME 0.277's `src/mame/sharp/x68k_hdc.cpp`
builds the SASI LBA from a 6-byte Group-0 CDB as
`(cmd[1] & 0x1f) << 16 | cmd[2] << 8 | cmd[3]` — **21 bits of 256-byte blocks,
so 512 MiB is addressable per unit**. `call_create` makes a 20 MB image
(`0x13c98` blocks) because that is what a period drive was.
So the 40 MB ceiling is a **Human68k / IPL volume-format and period-drive
limit**, not something the SASI command set imposes. That distinction does not
rescue the profile: four units at 40 MB is 160 MiB, and 146.0 MiB of video
would consume essentially the entire SASI address space of the machine at the
*lowest* rate this codec has ever produced, leaving nothing for Human68k, the
player, or the game's own data.
*Scope: the 21-bit CDB and the 256-byte block are read out of MAME's
implementation. The 40 MB volume figure is the user's, and is consistent with
Human68k's SASI partitioning; it has not been measured here.*
### 32.3 The rate point may come back, under a different name
Dropping SASI removes an interface, not a bitrate, and the two are on different
axes — the profile axis has been I/O bandwidth only since FINDINGS 28.7. The
awkward part is that **capacity and bandwidth now pull in opposite directions**:
- the only period medium with room for 371.8 MiB (let alone 599.6) is
**CD-ROM** at 540-650 MB, and
- a **1x CD-ROM sustains ~150 KB/s**, which is *below* the surviving 280 KB/s
profile and much nearer the rate that was just retired.
A SCSI hard disk has the bandwidth but has to be large for the era at 372 MiB.
**The user's call was to ship `scsi` as the only profile now and settle the
medium when the pipe is measured** — the blocked disk benchmark
(`docs/BENCHMARK.md`) and the DMA-vs-PIO check of FINDINGS 29.5.
**Correction to the framing above, found after that call was made.** The
medium is less open than this section first presented it. FINDINGS 21.2 already
committed the deployment target to **SD-backed SCSI (BlueSCSI / SCSI2SD)**, in
session 2, and that is the premise the whole 488 KB/s constant rests on. On SD
there is no capacity problem at any rate this codec produces — an SD card is
gigabytes — and no seek tail either. So:
- **Capacity does not choose between the survivors.** It killed SASI, whose
40 MB ceiling is a Human68k volume-format limit that SD emulation does not
lift, and it does not bind on SD-backed SCSI at all.
- **CD-ROM is the one that capacity rules out**, not in. With spans the stream
is 487.1 KB/s = **650.1 MiB** for the whole game, past a CD's ~620 MiB
usable — and 487 KB/s is more than 3x a 1x CD-ROM's ~150 KB/s. A CD-ROM
delivery would mean giving up the span lever *and* re-deriving a profile
around 150 KB/s.
So the open question is not "which medium" but the one FINDINGS 29.5/30.7
already had: **confirm the 488 KB/s figure's provenance, and confirm DMA**.
`profile_gen.py` exists precisely to re-derive a profile from a measured
bandwidth once there is one.
### 32.4 What MAME says about the SCSI path that survives
Read out of MAME 0.277 while settling 32.2, and directly relevant because the
medium decision is now the thing gating the profile:
**The CZ-6BS1's DMA is real and fully modelled.** `x68k_scsiext.cpp` wires the
MB89352's DREQ to the expansion slot and replaces the data register at
`$EA0015` with DMA-aware glue: on a DMA cycle (`m_slot->exown()`, driven by
`m_hd63450->own()`) a read goes to `spc->dma_r()` and #DTACK is negated until
DRQ asserts. `x68k.cpp:1114` closes the loop the other way
(`out_dtack_callback -> hd63450_device::dtack_w`). That is a genuine
DMAC-driven transfer with hardware flow control, on the **stock `x68000`
driver** — the one MAME marks working. This is the configuration FINDINGS 29.5
asked about, and the answer for this board is **DMA, not PIO**.
**The internal SCSI of the Super/XVI/030 is NOT modelled that way**, and it is
a trap. `x68k.cpp:1176` reads, verbatim, `// TODO: duplicate DMA glue from
CZ-6BS1`. So MAME's internal SCSI is PIO-only. A benchmark run on `x68ksupr`
would measure a PIO fallback the real machine does not have — on top of those
drivers already being MACHINE_NOT_WORKING (FINDINGS 28.7). **Benchmark
`x68000 -exp1 cz6bs1`, not `x68ksupr`.**
**CD-ROM is a first-class SCSI device on the internal bus**`x68k.cpp:1168`
puts an `NSCSI_CDROM` at ID 6 by default — but the CZ-6BS1 card's own device
list offers `harddisk` only. So the CD-ROM delivery route of 32.3 is
emulatable, but not on the board whose DMA is modelled, without a source
change.
None of this is a transfer RATE. `docs/BENCHMARK.md`'s split still holds and is
worth restating because 32.3 defers a decision to a measurement: MAME can
settle whether the path works and whether it is DMA, and **cannot** settle
KB/s, because its device models are functional rather than transfer-timing
accurate. The rate half of the medium question needs derivation or real
hardware, not a longer MAME run.
## 33. The container carries its own alignment: DLX1 -> DLX2 (session 9)
The encoder gap left open since session 7 (FINDINGS 28.3, STATUS item 4) is
closed. `encode.py` now emits **DLX2**, which pads every frame record up to a
4-byte boundary — the first one included, by padding the codebook tables so
`off_frm` is aligned. `dlx.py` reads both versions; DLX1 containers stay
readable because every measurement in FINDINGS 28-31 was taken on one.
Measured on the same 120-frame window:
| | record starts not 4-aligned | padding cost |
|---|---:|---:|
| DLX1 (through session 8) | **94/120** | 0 (the loader added 180 B) |
| DLX2 (now) | **0/120** | 160 B = 1.33 B/frame = **16 B/s** |
16 B/s against 278.6 KB/s is 0.006% of the stream. The thing it buys is not
speed: an odd `move.l (a0)+` on a 68000 is an **address error**, which vectors
into the IPL and presents as an infinite loop, not as a slow read. That is the
bug that cost session 7 an afternoon.
`tools/bench/prep_dlx.py` still realigns at load time and now says whether it
had to — `0/120 record starts unaligned -- the container carries its own
padding` on a DLX2 input. It is kept rather than deleted because it is what
makes the session 7-8 containers decodable, and those are the containers the
published timings belong to.
**Cross-check that this changed nothing else:** re-encoding the `scsi` window
with the DLX2 writer reproduces FINDINGS 31.1 exactly — 29.27 dB, 278.6 KB/s,
median 99.6% / max 110.6% of a 12fps frame, 1/120 frames missing. The padding
is additive; it does not touch the mode decision.
## 34. The cost model, checked against the machine on a cost-aware container (session 9)
STATUS item 1. Everything in FINDINGS 31 was the validated cost MODEL
(`vq_hybrid.cycles`) applied to a container it had never been checked against —
the 1-point validation of 28.2 belongs to the *session 7* stream. This is the
cost-aware container timed on the emulated 68000, same harness, same scope
(instruction cycles, zero-wait-state GVRAM, interrupts masked; a LOWER BOUND).
| anchor | non-SKIP | model | measured | error |
|---|---:|---:|---:|---:|
| min non-SKIP | 15.4% | 254,683 cyc / 30.6% | 262,751 / 31.5% | **-3.07%** |
| median | 53.2% | 681,199 / 81.7% | 690,251 / 82.8% | **-1.31%** |
| p90 | 75.7% | 832,116 / 99.9% | 834,213 / 100.1% | **-0.25%** |
| max non-SKIP | 100.0% | 921,293 / 110.6% | 921,187 / 110.5% | **+0.01%** |
| whole 120-frame mean | — | 649,089 / 77.9% | 657,081 / 78.8% | **-1.22%** |
The model holds, and its error is **signed**: it under-predicts by 1-3% on light
frames and converges to exact on heavy ones. That is the right direction to be
wrong in for a ceiling controller — the bisection is tightest where the model is
most accurate — but it means the median frame is ~1 point dearer than FINDINGS
31 reports, not cheaper.
The four synthetic single-mode frames reproduce session 7 exactly: all-V1
110.5%, all-V4 165.2%, all-RAW 147.6%, all-SKIP 4.9%. Those are properties of
`decode.s`, not of the container, so agreeing across two different streams is
the cross-check that the harness is measuring what it claims.
### 34.1 The 23-minute "hang" was the buffering trap again
The session-8 note said this run "was still going at 12 minutes of CPU". It was
re-run here and sat at 99.9% CPU for **23 minutes** with a 0-byte log, then was
killed. Re-launched under **`stdbuf -oL`** with `-seconds_to_run 60`, the
identical plan completed in about **25 seconds of wall time** and printed every
line as it went — MAME reports `Average speed: 528.72% (52 seconds)`, so the
whole plan needs ~52 emulated seconds and the machine runs it at 5x realtime.
The lesson is the one already in STATUS, one level deeper: it is not enough to
write MAME's output to a file instead of a pipe. **A file is block-buffered too**,
so a long MAME run is unobservable until it exits, and an unobservable run that
is merely finishing looks exactly like one that is wedged. Session 8 lost the
measurement to that, and session 9 lost 23 minutes to it before spending 25
seconds getting the answer. **`stdbuf -oL` on every MAME job that prints
progress.**
## 35. The CPU budget has never had the disk in it (session 9)
**Raised by the user: "PIO is such a CPU killer. DMA is not. I'm concerned about
us drawing the wrong conclusions."** The concern is correct, and it is larger
than the labelling question of 32.4. This is the seventh false premise this
project has caught, and the most expensive one.
Every CPU figure in FINDINGS 24 through 34 is measured against **833,333 cycles
per frame**, the full 10 MHz clock divided by 12 fps. Nothing has ever been
subtracted from it for moving the bitstream off the disk. The decoder has been
scored as though the data arrives for free.
### 35.1 What the transfer actually costs
`profile_gen.py` has carried `DMA_CLOCKS_PER_WORD = 8` since session 2 (FINDINGS
5, an **ESTIMATE** from HD63450 timing, never measured) and prints a "DMA steal"
line — but that line was only ever compared against the 38.3% *blit* figure of
FINDINGS 17, which FINDINGS 24 superseded and which was never the decoder cost.
It was never debited from the decoder budget.
At the rates that matter, on a 10 MHz 68000:
| stream | DMA @ 8 clk/word | PIO, unrolled (~12 clk/B) | PIO, byte loop (~20 clk/B) |
|---|---:|---:|---:|
| `scsi`, 278.6 KB/s | **11.4%** | 34.2% | 57.1% |
| `scsi` + spans, 487.1 KB/s | **20.0%** | 59.9% | **99.8%** |
The PIO columns are hand-derived floors, not measurements: a byte from an I/O
register plus a store is 16 cycles on a 68000 before any loop overhead. They are
here to size the risk, and the size of the risk is that **PIO at the span rate
consumes the entire machine**.
### 35.2 What that does to the conclusions of FINDINGS 31
Debiting the DMA steal — the *cheap* case, the one we are hoping for:
| | KB/s | steal | budget left | median | p90 | worst | fits? |
|---|---:|---:|---:|---:|---:|---:|:--|
| `scsi` today | 278.6 | 11.4% | 738,238 | **112.4%** | 112.9% | 124.8% | **no** |
| `scsi` + spans | 487.1 | 20.0% | 667,070 | **98.3%** | 102.8% | 114.3% | **no** |
FINDINGS 31's headline — "1 frame of 120 misses" — is measured against a budget
with no I/O in it. With DMA debited the surviving profile does not fit at all:
the *median* frame is over. And 31.6's "with spans the window fits completely"
becomes a worst frame of 114.3%, because **the span lever buys cycles by
spending bandwidth, and the bandwidth comes back out of the CPU as steal.**
Spans still help — 112.4% -> 98.3% at the median, 14 points — but they no longer
close the gap on their own.
### 35.3 Why this is not settled by the DMA finding of 32.4
32.4 established that the CZ-6BS1's DMA *path* exists and is modelled. Three
things it does not establish, and all three are load-bearing:
1. **DMA vs PIO is a property of OUR player, not of the board.** The hardware
supports DMA; if the player reads through IOCS and IOCS does PIO, we get PIO
and the table above. `docs/BENCHMARK.md` item 4 already proposed driving the
MB89352 registers directly for exactly this reason — that is now not an
optimisation but the difference between fitting and not.
2. **8 clocks per word has never been measured.** It is now the single most
load-bearing unmeasured number in the project: at 8 the port is marginal, at
12 it is dead, at 4 it is comfortable. It comes from a datasheet reading in
session 2 and nothing has checked it since.
3. **MAME cannot settle it.** Its device models are functional, not
transfer-timing accurate (BENCHMARK.md), and it models no GVRAM wait states
either — so a MAME run can confirm the transfer is a DMA cycle and cannot
price it. This needs derivation from the HD63450 and MB89352 datasheets, or
real hardware.
### 35.4 What this does and does not overturn
It does **not** overturn the decoder measurements: 300/448/400 cycles per block
and the model validation of FINDINGS 34 are properties of `decode.s` and stand
unchanged. What it overturns is every statement of the form "N frames of 120
miss the budget", because the budget was wrong. Those all need re-running
against `833,333 * (1 - steal)` once `steal` is a measurement rather than a
datasheet estimate.
It also sharpens the framerate question of STATUS item 5 considerably. At 10 fps
the budget is 1,000,000 cycles and the same DMA steal is proportionally smaller
per frame, which is now a much stronger argument for 10 fps than "one late frame
per cut" ever was.
### 35.5 `11_cpu_budget.py` now debits it, and 10 fps absorbs it
The tool takes `--io dma|pio|none` (default **dma**) and prints the budget it is
actually scoring against. On `tmp/rc_fr_singe_scsi_cpufit.dlx`:
| `--io` | budget left | median | worst | frames missing |
|---|---:|---:|---:|---:|
| `none` — the pre-session-9 premise | 833,333 | 99.6% | 110.6% | **1/120** |
| `dma` (8 clk/word, estimated) | 738,234 | 112.4% | 124.8% | **84/120** |
| `pio` (12 clk/B, floor) | 548,036 | 151.4% | 168.1% | **120/120** |
`--io none` prints a warning naming FINDINGS 35, so the old number cannot be
produced by accident.
**At 10 fps and DMA the same container goes back to 1/120** — median 93.7%,
worst 104.0%. That is conservative, because it holds the 12 fps byte rate: a
real 10 fps encode carries ~17% fewer bytes per second, so the steal falls too.
This changes what the framerate decision (STATUS item 5) is *for*. It was a
quality question about one late frame per scene cut. It is now the lever that
pays for the disk, and on current estimates it is the difference between a
stream that fits and one that misses 70% of its frames.
## 36. A `scsi` window does not fit in the machine the test rig emulates (session 9)
Swapping the decoder gate onto the surviving profile's container made it fail —
`frame 119 not pixel-exact: 49,005 px differ`. That is not a decoder bug and not
the DLX2 change. **The container does not fit in RAM.**
`tools/bench/decode.lua` loads the entire stream into emulated memory at
`STREAM = 0x30000`, and the locked target is a stock **2 MB** machine:
| container | stream | ends at | verdict |
|---|---:|---:|---|
| session 7-8 `sasi` | 1,108,888 B | 0x13EB98 = 1.25 MB | fits |
| `scsi` cost-aware | 2,840,860 B | 0x2E591C = 2.90 MB | **overruns 0x200000 by 940 KB** |
The loader wrote 940 KB past the top of memory, the decoder then parsed whatever
that reads back as, and the run neither completed its sequential pass nor drew
the right picture. Every 68000 decode verification before session 9 was done on
a container small enough to fit by accident — the `sasi` profile was a third the
bitrate, so nobody met this.
**This is a property of the test rig, not of the player.** The shipping player
streams from disk into a ring buffer and holds seconds of video, not minutes.
But it does bound what the rig can prove: at 278.6 KB/s, a 2 MB machine holds
about **6.7 seconds** of stream, so the strongest test in the tree can only ever
audit a prefix of a window.
The fix keeps the test honest rather than making it pass:
- `prep_dlx.py` truncates the frame list to what fits, **prints that it did**,
and takes `--ram` / `--all-frames`. On the `scsi` window it keeps **80 of
120** frames.
- `verify_decode.py` takes `--nframes` so the reference decoder replays exactly
the prefix the 68000 decoded, instead of running 40 frames ahead of it.
- `check.sh` reads the count back out of `decode_meta.lua` and passes it through,
and now **fails loudly if the sequential pass did not complete** — the missing
`snapshot taken` marker — instead of reporting a pixel diff against a
half-drawn screen. That guard is what turned this from a mystery into a
five-minute diagnosis.
Verifying a prefix is still a real test: SKIP blocks make every frame a claim
about the one before it, so frame 79 is only correct if all 80 were. What is
lost is coverage of the last 40 frames, and the honest way to get it back is to
gate on more than one window rather than to pretend one pass covers everything.
**The timing confirms the diagnosis.** Truncated to 80 frames the pass completes
in **8 emulated seconds** and the frame is pixel-exact; the model predicts
~6.6 s for 80 frames at this container's cost, so that is the expected number.
The 120-frame run that overran RAM could not finish the same work in **44**.
A decoder reading garbage does not run slowly for an interesting reason — it was
parsing lengths out of unmapped memory and walking wherever they pointed. Any
"the decoder is 4x slower than the model on RAW-heavy streams" conclusion drawn
from that run would have been entirely false, which is the third time in this
session that an unobservable run nearly produced a wrong finding.
+106 -75
View File
@@ -1,93 +1,112 @@
# Status & next-session handoff — end of session 8 (2026-08-23) # Status & next-session handoff — end of session 9 (2026-08-23)
## Where this stands ## Where this stands
Session 7 left the decoder pixel-exact and **31% of frames over the CPU budget** Session 9 did three things: dropped a profile on the user's instruction, closed
at `sasi`, 42% at `scsi` (FINDINGS 28), with two levers proposed and neither the last encoder gap, validated the cost model against the machine — and then
measured. Session 8 did both. found that **the CPU budget every one of those numbers was scored against has
never had the disk in it**.
**Lever A, spans: measured.** A row-linear span of word-expanded literals costs **FINDINGS 35 is the headline, and it is bad news.** Raised by the user ("PIO is
**43.7 cycles per span + 9.152 per pixel** — but only in an encoder-assisted such a CPU killer. DMA is not. I'm concerned about us drawing the wrong
format, `{u32 absolute GVRAM address, u16 jump displacement}` into an unrolled conclusions"). Every CPU figure in FINDINGS 24-34 is measured against the full
copy chain. The obvious decoder, handed `(x, npix)`, is 97.9 + 10.46. 833,333 cycles/frame, with nothing subtracted for moving the bitstream off SCSI.
FINDINGS 30, `tools/bench/span.sh` (~25 s). Debiting the HD63450 cycle-steal at the long-standing 8 clocks/word estimate:
**Lever B, the cost-aware mode decision: implemented, measured, and ON by | `scsi` container | budget left | median frame | worst | frames missing |
default.** `decide()` minimises `D + lam*bytes + mu*cycles`; `mu` is bisected |---|---:|---:|---:|---:|
per frame against 833,333 cycles with the `lam` bisection nested inside it. | no I/O — the pre-session-9 premise | 833,333 | 99.6% | 110.6% | **1/120** |
FINDINGS 31, `tools/analysis/13_cpu_ratectl.py`. | **DMA** (8 clk/word, ESTIMATED) | 738,234 | **112.4%** | 124.8% | **84/120** |
| PIO (12 clk/B, hand-derived floor) | 548,036 | 151.4% | 168.1% | **120/120** |
| `sasi`, worst sustained window | PSNR | KB/s | CPU median | CPU max | missing | So "1 frame of 120 misses" was against the wrong budget. `11_cpu_budget.py` now
|---|---:|---:|---:|---:|---:| takes `--io dma|pio|none`, defaults to `dma`, and warns if asked for `none`.
| bytes only (session 7) | 27.22 dB | 109.5 | 74.4% | 136.2% | **37/120** |
| + cycle ceiling (now the default) | 26.95 dB | 109.4 | 81.5% | 110.6% | **1/120** |
| + ceiling + spans (MODEL, nothing implements it) | — | 449.3 | 56.8% | 91.5% | **0/120** |
`scsi`: 51/120 -> 1/120 for 0.62 dB. The one remaining miss at either profile is **The way out is 10 fps, and it works on paper**: same container, DMA debited,
**frame 0**, which has no previous reconstruction and so is 100% changed by 10 fps -> median 93.7%, worst 104.0%, **1/120** — and that is conservative,
definition — the same case as a scene cut. It is emitted late on purpose. because it holds the 12 fps byte rate. FINDINGS 35.5.
The cost-aware container is verified pixel-exact on the 68000 (120 frames, **SASI is dropped (USER DECISION).** A SASI volume is 40 MB and the game's 22.8
`tools/bench/verify_decode.py`). minutes of unique footage is 146 MiB at the *lowest* rate this codec makes.
`scsi` is the only profile. Delivery is SD-backed SCSI, as locked in session 2 —
capacity does not bind there, and it is what rules CD-ROM out. FINDINGS 32.
**The encoder gap is closed.** `encode.py` emits **DLX2**, 4-byte-aligned frame
records: 94/120 record starts were on odd addresses, now 0/120, for 16 B/s.
Re-encoding reproduces FINDINGS 31.1 exactly. FINDINGS 33.
**The cost model is validated on a cost-aware container** (STATUS item 1, done):
errors -3.07% to +0.01% against the 68000, whole-window mean -1.22%. It
under-predicts light frames and is exact on heavy ones. FINDINGS 34.
**A `scsi` window does not fit the 2 MB machine the rig emulates** — 2.84 MB of
stream against a 0x200000 ceiling. The decoder gate now verifies 80 of 120
frames and says so, rather than failing as a phantom pixel diff. FINDINGS 36.
## NEXT SESSION, in order ## NEXT SESSION, in order
0. **Green light first.** `./tools/bench/check.sh` (~3 min, Blu-ray mounted). 0. **Green light first.** `./tools/bench/check.sh` (~5 min, Blu-ray mounted).
The drift stage now runs BOTH controllers; both must report 0/120. Verified green at end of session 9. The gate container is now
`tmp/rc_fr_singe_scsi_cpufit.dlx`, and it decodes **80 of 120 frames**, not
120: the rig loads the whole stream into a 2 MB machine and a `scsi` window
is 2.84 MB, which overran the top of RAM by 940 KB and produced a failure
that looked exactly like a decoder bug. `prep_dlx.py` now truncates to what
fits and says so; `verify_decode.py --nframes` replays the same prefix; and
the stage fails loudly if the sequential pass did not complete. FINDINGS 36.
1. **Time the cost-aware container on the 68000.** Everything in FINDINGS 31 is At 278.6 KB/s a 2 MB machine holds ~6.7 s of stream, so **the strongest test
the validated cost MODEL (`vq_hybrid.cycles`, within 1 point of the machine in the tree can only ever audit a prefix of a window.** Getting the coverage
on four frames of the session-7 container), not a measurement of this one. back means gating on more than one window, not one longer pass.
The full timing pass was started and killed for time:
```
python3 tools/bench/prep_dlx.py tmp/rc_fr_singe_sasi_cpufit.dlx
tools/vasm/vasmm68k_mot -Fbin -o tmp/decode.bin src/player/decode.s
( cd tmp && SDL_VIDEODRIVER=dummy mame x68000 -bios ipl10 -ramsize 2M \
-video soft -window -sound none -nothrottle -plugins \
-autoboot_script ../tools/bench/decode.lua -snapshot_directory ./snap_decode \
-snapview native -seconds_to_run 300 > decode_cpufit.log 2>&1 )
```
**Budget real time for it: the run was still going at 12 minutes of CPU.**
MAME's stdout is block-buffered to a file, so there is no progress to watch
— wait on the PID, never on a `pgrep -f` match (see the shell traps below).
Confirm the four anchors against `11_cpu_budget.py` on the same container,
and update FINDINGS 31 with measured-vs-model errors.
2. **Put spans in the bitstream.** This is the big one and it is now fully 1. **Measure the DMA cycle-steal, or derive it from primary sources.** This is
specified by measurement: format in FINDINGS 30.2, costs in 30.5, and the now the most load-bearing unmeasured number in the project: at 8 clocks/word
scene-cut arithmetic in 30.6. It touches `encode.py` (a fifth mode and a the port is marginal, at 12 it is dead, at 4 it is comfortable. It has been
run-aware decision), `dlx.py` (the reference decoder), and `decode.s`. The an estimate from a session-2 datasheet reading since FINDINGS 5. MAME cannot
24-pixel quantisation and the free row overrun are part of the format, not answer it (functional models, not transfer-timing accurate) — this needs the
optimisations to add later. Order it AFTER item 1 so the model that prices it HD63450 and MB89352 datasheets or real hardware. **Everything below is
has been checked against the machine once more. contingent on it.**
3. **The three open items of FINDINGS 29.5/30.7**, now load-bearing because a 2. **Make sure the player actually gets DMA.** DMA-vs-PIO is a property of our
span design runs at ~449 KB/s of a 488 KB/s pipe: re-run the ring-buffer code, not the board: the CZ-6BS1's DMA path is real and modelled
simulation at that rate (FINDINGS 21 was established at 110 and 280), confirm (FINDINGS 32.4), but if the player reads through IOCS and IOCS does PIO we
the provenance of the user's 4 Mbps figure, and **confirm DMA rather than get the 120/120 row. `docs/BENCHMARK.md` item 4 (drive the MB89352 directly)
PIO** — a PIO fallback puts a 449 KB/s transfer back on the CPU the whole is no longer an optimisation. Benchmark `x68000 -exp1 cz6bs1`, **never
lever exists to relieve. The DMA check is the cheapest of the three and the `x68ksupr`** — MAME's internal SCSI has no DMA glue at all (`// TODO:
most consequential. duplicate DMA glue from CZ-6BS1`) and would measure a PIO fallback the real
machine does not have.
4. **Encoder gap, still open from session 7:** `encode.py` should pad frame 3. **Re-decide the framerate, now that it is the lever that pays for the disk.**
records to 4 bytes. Frame boundaries land on odd addresses and a 68000 takes 10 fps absorbs the DMA steal on current estimates. This was item 5 and a
an address error, not a slow read (FINDINGS 28.3). `prep_dlx.py` pads at load quality question; it is now arithmetic. It is still the user's call.
time, which is why the decoder works; the container itself does not.
Measured cost of fixing it: 1.5 B/frame = 18 B/s.
5. **A quality-vs-framerate question that is the user's, not the encoder's.** 4. **Then spans.** Fully specified by measurement (format FINDINGS 30.2, costs
Every miss is now one frame per cut. The options remain: one late frame at 30.5, scene-cut arithmetic 30.6). Note what 35.2 does to the case for them:
each cut (the outgoing content is unrelated, so it may be invisible), a cut spans buy cycles by spending bandwidth, and the bandwidth comes back as
spread over two frame times, or 10fps. Spans (item 2) make the question go steal, so 31.6's "fits completely" becomes a worst frame of 114.3%. They are
away if they land as modelled. still worth 14 points at the median — they are just no longer sufficient
alone.
5. **Re-run the ring-buffer simulation at the surviving rate** and confirm the
488 KB/s figure's provenance (FINDINGS 29.5/30.7, still open).
**Do not start by hand-optimising `decode.s`.** The hand-derived timings agree **Do not start by hand-optimising `decode.s`.** The hand-derived timings agree
with the measurements to 0.5% on V1 and 1% on RAW (FINDINGS 28.4), so the inner with the measurements to 0.5% on V1 and 1% on RAW (FINDINGS 28.4), and FINDINGS
loop is close to what the instruction set allows. 34 confirms the model on a second container. The inner loop is close to what the
instruction set allows; the cycles to be won are in the budget, not the loop.
**Three of this session's near-misses had the same shape: an unobservable run
almost produced a false finding.** The 23-minute "hang" (FINDINGS 34.1), the
RAM overrun that looked like a 4x-slow decoder (FINDINGS 36), and the truncated
gate that reported 49,005 differing pixels. In each case the instrument was
broken, not the thing being measured. **Always `stdbuf -oL` a MAME job that
prints progress.** A file is
block-buffered too, so a long run is unobservable until it exits — and an
unobservable run that is merely finishing looks exactly like one that is wedged.
Session 8 lost this measurement to that; session 9 lost 23 minutes to it before
spending 25 seconds getting the answer with `stdbuf`. FINDINGS 34.1.
--- ---
## What session 8 settled ## What session 8 settled
0. **The mode decision can see cycles, it is on by default, and it costs 0. **The mode decision can see cycles, it is on by default, and it costs
@@ -245,22 +264,34 @@ The two session-7 stages are worth knowing the shape of before they fail on you:
| Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c | | Display mode | 256 colors, 256x192 in 256x256 CRTC mode | every mode is 1 word-access/pixel, so 256c is free vs 16c |
| Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway | | Double buffer | **none** — page 1 sacrificed | enables `movem.l` 24px bursts; delta coding needs a RAM reference frame anyway |
| **Codec** | **hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion** | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 | | **Codec** | **hybrid VQ: SKIP / V1 4x4 / V4 four-2x2 / RAW, per-block rate-distortion** | flat 4x4 VQ was measured and rejected — see FINDINGS 9-10 |
| **Quality modes** | **two: `sasi` and `scsi`** (USER DECISION, session 2) | one codec, one decoder, one bitstream; only `lam` differs | | **Quality modes** | **one: `scsi`** (USER DECISION, session 9 — `sasi` dropped) | a SASI volume is 40 MB and the game is 146 MiB at the LOWEST rate this codec makes. FINDINGS 32 |
| Delivery medium | **SD-backed SCSI** (BlueSCSI / SCSI2SD), as locked in session 2 | capacity does not bind on SD at any rate this codec makes; it is what killed SASI, and it is what rules CD-ROM OUT (650 MiB with spans, and 1x CD is ~150 KB/s). FINDINGS 32.3 |
| Profile axis | **I/O bandwidth only** | the profiles say nothing about CPU; **both target the same stock 10MHz 68000**, and the Super has SCSI at 10MHz. FINDINGS 28.7 | | Profile axis | **I/O bandwidth only** | the profiles say nothing about CPU; **both target the same stock 10MHz 68000**, and the Super has SCSI at 10MHz. FINDINGS 28.7 |
| Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win | | Framerate | 12 fps, **explicit decimation** | source has zero duplicate frames; no free "twos" win |
| Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something | | Emulator | MAME 0.277 x68000 | accurate enough that measured cycles mean something |
| SNES project reuse | **MIT — cleared** | `data/events/` scene graph is reusable with attribution | | SNES project reuse | **MIT — cleared** | `data/events/` scene graph is reusable with attribution |
### The SASI/SCSI question is RESOLVED ### The SASI/SCSI question is RESOLVED — and in session 9 SASI was DROPPED
Session 1 left "which machine do we target" open. The user's answer: **ship both**, Session 1 left "which machine do we target" open. Session 2's answer was **ship
as two quality profiles. This is now implemented rather than hypothetical — the both**, as two quality profiles. **Session 9 retired `sasi`** (USER DECISION) on
bitrate ceiling is a build parameter in `tools/encoder/ratectl.py`: CAPACITY, not bandwidth: a SASI volume is limited to 40 MB, and the 22.8 minutes
of unique scene footage on the source Blu-ray is **146 MiB even at 110 KB/s** —
more than the machine's whole 4-unit SASI address space. FINDINGS 32.
Everything below this line about *two* profiles is the session-2..8 record, kept
because every measurement in FINDINGS 27-31 was taken against it. One profile
ships:
| profile | target | lam | quality (00020 / 00146) | machine | | profile | target | lam | quality (00020 / 00146) | machine |
|---|---|---|---|---| |---|---|---|---|---|
| `sasi` | 110 KB/s | 60 (floor) | 36.9 / 29.6 dB | stock 10MHz ACE/EXPERT | | ~~`sasi`~~ | ~~110 KB/s~~ | ~~60~~ | ~~36.9 / 29.6 dB~~ | **RETIRED session 9** |
| `scsi` | 280 KB/s | 10 (floor) | 39.4 / 32.3 dB | Super/XVI, or CZ-6BS1 board | | `scsi` | 280 KB/s | 10 (floor) | 39.4 / 32.3 dB | Super/XVI, or CZ-6BS1 board |
The 110 KB/s **rate point** is not necessarily gone with the interface: a 1x
CD-ROM sustains ~150 KB/s and CD-ROM is the only period medium with room for the
whole game (374 MiB at `scsi`, 600 MiB with spans). The user's call was to ship
one profile now and settle the medium when the pipe is measured. FINDINGS 32.3.
**That "machine" column is about the BUS, not the CPU.** The profiles are an **That "machine" column is about the BUS, not the CPU.** The profiles are an
I/O-bandwidth axis and say nothing about clock speed: the X68000 Super has I/O-bandwidth axis and say nothing about clock speed: the X68000 Super has
built-in SCSI at 10 MHz (`x68k.cpp:1194`, `40_MHz_XTAL/4`, same as the base built-in SCSI at 10 MHz (`x68k.cpp:1194`, `40_MHz_XTAL/4`, same as the base
+2 -2
View File
@@ -16,7 +16,7 @@ the mode headers would exploit.
RAW 16 literal palette indices -- the escape that makes lam=0 pixel-exact RAW 16 literal palette indices -- the escape that makes lam=0 pixel-exact
Usage: python3 tools/analysis/08_mode_map.py <frames_dir> <out.webm> Usage: python3 tools/analysis/08_mode_map.py <frames_dir> <out.webm>
[--profile sasi|scsi] [--scale N] [--lossless] [--fixed-lam] [--profile scsi] [--scale N] [--lossless] [--fixed-lam]
--fixed-lam renders the pre-session-6 encoder (no rate control) instead. --fixed-lam renders the pre-session-6 encoder (no rate control) instead.
Output format follows the extension. Prefer .webm: GIF re-quantises to 256 Output format follows the extension. Prefer .webm: GIF re-quantises to 256
@@ -45,7 +45,7 @@ def main():
if "--scale" in sys.argv: if "--scale" in sys.argv:
SCALE = int(sys.argv[sys.argv.index("--scale")+1]) SCALE = int(sys.argv[sys.argv.index("--scale")+1])
prof = RC.PROFILES[sys.argv[sys.argv.index("--profile")+1] prof = RC.PROFILES[sys.argv[sys.argv.index("--profile")+1]
if "--profile" in sys.argv else "sasi"] if "--profile" in sys.argv else "scsi"]
m = H.build(src, k1=prof["k1"], k4=prof["k4"]) m = H.build(src, k1=prof["k1"], k4=prof["k4"])
# Rate-controlled by default, so the map shows the mode decisions that # Rate-controlled by default, so the map shows the mode decisions that
# actually ship. --fixed-lam renders the pre-session-6 encoder instead; # actually ship. --fixed-lam renders the pre-session-6 encoder instead;
+44 -5
View File
@@ -23,6 +23,8 @@ sys.path.insert(0, "tools/encoder")
import numpy as np import numpy as np
from dlx import DLX from dlx import DLX
import vq_hybrid as H import vq_hybrid as H
import ratectl as RC
RC_AUDIO_BPS = RC.AUDIO_KBPS * 1024
# Machine clocks, confirmed from MAME 0.277 src/mame/sharp/x68k.cpp:1133/1194/ # Machine clocks, confirmed from MAME 0.277 src/mame/sharp/x68k.cpp:1133/1194/
# 1200 -- not recalled. x68000 and x68ksupr are BOTH 40_MHz_XTAL/4 = 10 MHz; # 1200 -- not recalled. x68000 and x68ksupr are BOTH 40_MHz_XTAL/4 = 10 MHz;
@@ -45,6 +47,19 @@ ap.add_argument("container", nargs="?",
ap.add_argument("--machine", default="stock", choices=list(CLOCKS), ap.add_argument("--machine", default="stock", choices=list(CLOCKS),
help="which X68000's clock to budget against (default stock)") help="which X68000's clock to budget against (default stock)")
ap.add_argument("--fps", type=float, default=FPS) ap.add_argument("--fps", type=float, default=FPS)
# FINDINGS 35: the frame budget has never had the disk in it. The bitstream has
# to be moved off SCSI into the ring buffer, and on this machine that costs CPU
# whether it is DMA (the HD63450 cycle-steals) or PIO (the 68000 moves every
# byte). Default ON, because scoring a decoder against a budget that assumes the
# data arrives for free is exactly the mistake 35 was raised to stop.
ap.add_argument("--io", default="dma", choices=["dma", "pio", "none"],
help="how the bitstream reaches RAM (default dma)")
ap.add_argument("--dma-clocks-per-word", type=float, default=8.0,
help="HD63450 cycle-steal. ESTIMATE from FINDINGS 5, NEVER "
"MEASURED, and the most load-bearing unmeasured number "
"in the project (FINDINGS 35.3)")
ap.add_argument("--pio-clocks-per-byte", type=float, default=12.0,
help="hand-derived floor for a 68000 register-to-RAM copy")
a = ap.parse_args() a = ap.parse_args()
CPUHZ = CLOCKS[a.machine] * 1e6 CPUHZ = CLOCKS[a.machine] * 1e6
FPS = a.fps FPS = a.fps
@@ -54,14 +69,35 @@ if not os.path.exists(a.container):
d = DLX(a.container) d = DLX(a.container)
# --- what the transfer costs, from the container's own byte rate
vid_bps = sum(n + 4 for (_, n) in d.frames) / d.nframes * d.fps
io_bps = vid_bps + RC_AUDIO_BPS
if a.io == "dma":
io_cycles_per_s = (io_bps / 2) * a.dma_clocks_per_word
elif a.io == "pio":
io_cycles_per_s = io_bps * a.pio_clocks_per_byte
else:
io_cycles_per_s = 0.0
io_pct = 100 * io_cycles_per_s / CPUHZ
FRAME_NET = FRAME * (1 - io_pct / 100)
modes = [d.modes(f) for f in range(d.nframes)] modes = [d.modes(f) for f in range(d.nframes)]
cyc = np.array([cycles(m) for m in modes]) cyc = np.array([cycles(m) for m in modes])
pct = 100 * cyc / FRAME pct = 100 * cyc / FRAME_NET
ns = np.array([100 * (m != 0).mean() for m in modes]) ns = np.array([100 * (m != 0).mean() for m in modes])
print(f"{a.container}: {d.nframes} frames, {d.nb} blocks/frame") print(f"{a.container}: {d.nframes} frames, {d.nb} blocks/frame")
print(f"budget: {a.machine} @ {CLOCKS[a.machine]:.2f} MHz, {FPS:g} fps " print(f"budget: {a.machine} @ {CLOCKS[a.machine]:.2f} MHz, {FPS:g} fps "
f"-> {FRAME:,.0f} cycles/frame") f"-> {FRAME:,.0f} cycles/frame")
print(f" I/O ({a.io}): {io_bps/1024:.1f} KB/s costs {io_pct:.1f}% of the CPU "
f"-> {FRAME_NET:,.0f} cycles/frame left for decoding")
if a.io == "dma":
print(f" {a.dma_clocks_per_word:g} clocks/word is an ESTIMATE (FINDINGS 5), "
f"never measured -- see FINDINGS 35.3")
elif a.io == "none":
print(" WARNING: --io none scores the decoder as if the disk were free. "
"That is the\n premise FINDINGS 35 overturned; every 'N frames miss' "
"figure before session 9\n was computed this way.")
if a.machine != "stock": if a.machine != "stock":
print(" (derived: scaled by clock from cycles measured on the 10 MHz core.\n" print(" (derived: scaled by clock from cycles measured on the 10 MHz core.\n"
" MAME 0.277 marks x68ksupr/x68kxvi/x68030 MACHINE_NOT_WORKING, so\n" " MAME 0.277 marks x68ksupr/x68kxvi/x68030 MACHINE_NOT_WORKING, so\n"
@@ -100,10 +136,11 @@ if a.machine == "stock" and a.fps == 12:
f"(optimistic by {np.median(pct)/np.median(old):.2f}x at the median)") f"(optimistic by {np.median(pct)/np.median(old):.2f}x at the median)")
miss = pct > 100 miss = pct > 100
print(f"\nframes that do NOT fit {FRAME:,.0f} cycles: {miss.sum()}/{d.nframes} " print(f"\nframes that do NOT fit {FRAME_NET:,.0f} cycles: {miss.sum()}/{d.nframes} "
f"({100*miss.mean():.0f}%)") f"({100*miss.mean():.0f}%)")
print(f" sustainable framerate if EVERY frame must fit: " print(f" sustainable framerate if EVERY frame must fit: "
f"{CPUHZ/cyc.max():.1f} fps; at the mean frame {CPUHZ/cyc.mean():.1f} fps") f"{CPUHZ*(1-io_pct/100)/cyc.max():.1f} fps; at the mean frame "
f"{CPUHZ*(1-io_pct/100)/cyc.mean():.1f} fps")
if miss.any(): if miss.any():
print(f" worst {pct.max():.1f}% -- {(pct.max()-100)/100*1000/FPS:.0f} ms late " print(f" worst {pct.max():.1f}% -- {(pct.max()-100)/100*1000/FPS:.0f} ms late "
f"on an {1000/FPS:.0f} ms frame") f"on an {1000/FPS:.0f} ms frame")
@@ -116,5 +153,7 @@ print(f"\nwhere the cycles go, over the whole window:")
for k, n in enumerate(("SKIP", "V1", "V4", "RAW")): for k, n in enumerate(("SKIP", "V1", "V4", "RAW")):
print(f" {n:<5} {100*tot[k]/tot.sum():5.1f}% of blocks " print(f" {n:<5} {100*tot[k]/tot.sum():5.1f}% of blocks "
f"{100*spend[k]/spend.sum():5.1f}% of the cycles") f"{100*spend[k]/spend.sum():5.1f}% of the cycles")
print(f"\nV4 is {C_V4/C_V1:.2f}x a V1 block for {4}x the payload bytes -- the mode " print(f"\nV4 is {C_V4/C_V1:.2f}x a V1 block for {4}x the payload bytes. Since "
f"decision\nin vq_hybrid.py charges it the bytes but not the cycles.") f"session 8 the mode\ndecision charges it BOTH (decide(ctx, lam, mu), "
f"FINDINGS 31), which is why V4 is now\nthe rarest non-SKIP mode here -- "
f"a byte-rich profile buys its way out to RAW instead.")
+6 -4
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""What does fitting the CPU budget cost in quality? (session 8, lever B) """What does fitting the CPU budget cost in quality? (session 8, lever B)
python3 tools/analysis/13_cpu_ratectl.py [frames_dir] [--profiles sasi,scsi] python3 tools/analysis/13_cpu_ratectl.py [frames_dir] [--profiles scsi]
Session 6 made the BYTE budget a ceiling by bisecting `lam` per frame. FINDINGS Session 6 made the BYTE budget a ceiling by bisecting `lam` per frame. FINDINGS
28 then showed the binding budget is CYCLES, not bytes, and that the mode 28 then showed the binding budget is CYCLES, not bytes, and that the mode
@@ -28,7 +28,7 @@ import vq as VQ, vq_hybrid as H, ratectl as RC
ap = argparse.ArgumentParser() ap = argparse.ArgumentParser()
ap.add_argument("frames_dir", nargs="?", default="tmp/fr_singe") ap.add_argument("frames_dir", nargs="?", default="tmp/fr_singe")
ap.add_argument("--profiles", default="sasi,scsi") ap.add_argument("--profiles", default="scsi")
ap.add_argument("--fps", type=int, default=12) ap.add_argument("--fps", type=int, default=12)
ap.add_argument("--cache", default=None, help="pickle of H.build (auto by dir)") ap.add_argument("--cache", default=None, help="pickle of H.build (auto by dir)")
a = ap.parse_args() a = ap.parse_args()
@@ -97,5 +97,7 @@ for name in a.profiles.split(","):
print() print()
print("FINDINGS 28.7: re-coding every non-SKIP block as V1 is the floor the " print("FINDINGS 28.7: re-coding every non-SKIP block as V1 is the floor the "
"CURRENT mode set\nallows, and it still misses 11 frames at sasi / 12 at " "CURRENT mode set\nallows, and it still missed 11 frames at the retired "
"scsi. Misses above that floor\nare item 4 (spans), not item 1.") "110 KB/s profile / 12 at scsi.\nMisses above that floor are spans, not "
"the mode decision -- and 31.3 showed the\nfloor itself was too "
"pessimistic, because the real decision can move a block to SKIP.")
+28 -5
View File
@@ -57,17 +57,40 @@ echo "--- session 7: 68000 decoder is pixel-exact (FINDINGS 28) ---"
# 68000 code, every block mode, full temporal recursion. A SKIP block is a claim # 68000 code, every block mode, full temporal recursion. A SKIP block is a claim
# about the previous frame still being on screen, so the last frame is only # about the previous frame still being on screen, so the last frame is only
# right if all 120 were. # right if all 120 were.
DLX=tmp/rc_fr_singe_sasi_rcprofile.dlx # The gate container is the CURRENT default encode: scsi (the only profile left
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile sasi # after session 9 dropped sasi on capacity, FINDINGS 32), cost-aware mode
# decision on, DLX2 4-byte-aligned records. It is also the heavier stream --
# 43% RAW against sasi's 10% -- so it exercises the decoder harder than the
# session-7 container this gate used to run on.
DLX=tmp/rc_fr_singe_scsi_cpufit.dlx
[ -f "$DLX" ] || python3 tools/encoder/encode.py tmp/fr_singe "$DLX" --profile scsi
python3 tools/bench/prep_dlx.py "$DLX" > tmp/prep_dlx.log python3 tools/bench/prep_dlx.py "$DLX" > tmp/prep_dlx.log
# The rig loads the whole stream into a 2 MB machine, so a scsi window does not
# fit and prep_dlx truncates it. Verify against exactly the prefix it emitted.
NF=$(sed -n 's/.*nframes=\([0-9]*\),.*/\1/p' tmp/decode_meta.lua)
grep -a "TRUNCATED" tmp/prep_dlx.log || true
tools/vasm/vasmm68k_mot -Fbin -o tmp/decode.bin src/player/decode.s > /dev/null tools/vasm/vasmm68k_mot -Fbin -o tmp/decode.bin src/player/decode.s > /dev/null
mkdir -p tmp/snap_decode mkdir -p tmp/snap_decode
rm -f tmp/snap_decode/x68000/*.png rm -f tmp/snap_decode/x68000/*.png
( cd tmp && DLX_VERIFY_ONLY=1 SDL_VIDEODRIVER=dummy timeout -k 5 300 mame x68000 \ # stdbuf -oL: a FILE is block-buffered too, so without it a long MAME run is
# unobservable until it exits and a run that is merely finishing looks exactly
# like one that is wedged (FINDINGS 34.1).
# -seconds_to_run must cover the WHOLE sequential pass. The scsi container is
# 2.7x the payload of the session-7 one this gate used to run on, and at 20 s
# the pass was truncated -- MAME exited mid-decode and verify_decode.py then
# compared a partially drawn screen and reported 49,005 differing pixels, which
# reads as a decoder bug and is not one.
( cd tmp && DLX_VERIFY_ONLY=1 SDL_VIDEODRIVER=dummy stdbuf -oL timeout -k 5 300 mame x68000 \
-bios ipl10 -ramsize 2M -video soft -window -sound none -nothrottle -plugins \ -bios ipl10 -ramsize 2M -video soft -window -sound none -nothrottle -plugins \
-autoboot_script ../tools/bench/decode.lua \ -autoboot_script ../tools/bench/decode.lua \
-snapshot_directory ./snap_decode -snapview native -seconds_to_run 20 \ -snapshot_directory ./snap_decode -snapview native -seconds_to_run 45 \
> decode_check.log 2>&1 ) > decode_check.log 2>&1 )
python3 tools/bench/verify_decode.py "$DLX" # A truncated run must fail as a truncated run. Without this the only symptom is
# a pixel diff against a half-drawn frame.
grep -q "snapshot taken" tmp/decode_check.log || {
echo "FAIL: the 68000 sequential pass did not complete -- no snapshot marker."
echo " Raise -seconds_to_run; the pass needs the whole container decoded."
tail -5 tmp/decode_check.log; exit 1; }
python3 tools/bench/verify_decode.py "$DLX" --nframes "$NF"
echo "ALL GREEN" echo "ALL GREEN"
+41 -4
View File
@@ -29,9 +29,26 @@ sys.path.insert(0, "tools/encoder")
import numpy as np import numpy as np
from dlx import DLX from dlx import DLX
# The harness loads the WHOLE container into emulated RAM at STREAM=0x30000 and
# the target is a stock 2 MB machine, so there is a hard ceiling on how much of
# a stream can be verified in one pass. The shipping player streams from disk
# into a ring buffer and has no such limit; this is a property of the test rig.
# A `scsi` window overruns it -- 2.84 MB of stream ends at 0x2E591C, 940 KB past
# the 0x200000 top of RAM -- so the frame list is truncated to what fits and the
# truncation is announced. Verifying a prefix is still a real test: SKIP blocks
# make every frame a claim about the one before it.
STREAM_BASE = 0x30000
RAM_TOP = 0x200000
MARGIN = 0x8000 # stack, flags, codebooks live below STREAM_BASE
ap = argparse.ArgumentParser() ap = argparse.ArgumentParser()
ap.add_argument("container") ap.add_argument("container")
ap.add_argument("--out", default="tmp/decode") ap.add_argument("--out", default="tmp/decode")
ap.add_argument("--ram", type=lambda v: int(v, 0), default=RAM_TOP,
help="top of emulated RAM (default 0x200000, a stock 2 MB machine)")
ap.add_argument("--all-frames", action="store_true",
help="do NOT truncate to what fits in RAM (the loader will write "
"past the top of memory and the decoder will read garbage)")
a = ap.parse_args() a = ap.parse_args()
d = DLX(a.container) d = DLX(a.container)
@@ -65,12 +82,24 @@ dark = int(((render(I).astype(int)) ** 2).sum(1).argmin())
# so this loader realigns it; the encoder should carry the padding itself # so this loader realigns it; the encoder should carry the padding itself
# (FINDINGS 28.3). It costs at most 3 bytes per frame -- 36 B/s at 12fps, # (FINDINGS 28.3). It costs at most 3 bytes per frame -- 36 B/s at 12fps,
# against a 110 KB/s budget. # against a 110 KB/s budget.
budget = a.ram - STREAM_BASE - MARGIN
stream, rec_off, pad = bytearray(), [], 0 stream, rec_off, pad = bytearray(), [], 0
dropped = 0
for (o, n) in d.frames: for (o, n) in d.frames:
while len(stream) % 4: while len(stream) % 4:
stream += b"\0"; pad += 1 stream += b"\0"; pad += 1
if not a.all_frames and len(stream) + 4 + n > budget:
dropped = d.nframes - len(rec_off)
break
rec_off.append(len(stream)) rec_off.append(len(stream))
stream += n.to_bytes(4, "big") + d.raw[o:o + n] stream += n.to_bytes(4, "big") + d.raw[o:o + n]
NFRAMES = len(rec_off)
if dropped:
print(f" TRUNCATED: {NFRAMES}/{d.nframes} frames fit in RAM "
f"(stream budget {budget:,} B at 0x{STREAM_BASE:X} under a "
f"{a.ram/1024/1024:.0f} MB machine); {dropped} frames dropped.\n"
f" This is the TEST RIG's limit, not the player's -- the player "
f"streams into a ring buffer.")
# Synthetic single-mode frames. No real frame is all one mode, but the mix is # Synthetic single-mode frames. No real frame is all one mode, but the mix is
# exactly what the "76.6% x non-SKIP fraction" model of FINDINGS 24.5 assumes # exactly what the "76.6% x non-SKIP fraction" model of FINDINGS 24.5 assumes
@@ -86,7 +115,7 @@ for name, mo, per in (("all-SKIP", 0, 0), ("all-V1", 1, 1),
stream += (d.mode_bytes + d.nb * per).to_bytes(4, "big") + hdr + bytes(d.nb * per) stream += (d.mode_bytes + d.nb * per).to_bytes(4, "big") + hdr + bytes(d.nb * per)
# --- timing anchors: the distribution, not its mean (FINDINGS 25.6's lesson) # --- timing anchors: the distribution, not its mean (FINDINGS 25.6's lesson)
ns = np.array([100 * (d.modes(i) != 0).mean() for i in range(d.nframes)]) ns = np.array([100 * (d.modes(i) != 0).mean() for i in range(NFRAMES)])
order = np.argsort(ns) order = np.argsort(ns)
pick = { pick = {
"min non-SKIP %.1f%%" % ns[order[0]]: int(order[0]), "min non-SKIP %.1f%%" % ns[order[0]]: int(order[0]),
@@ -104,7 +133,7 @@ open(a.out + "_data.bin", "wb").write(blob)
with open(a.out + "_meta.lua", "w") as fh: with open(a.out + "_meta.lua", "w") as fh:
fh.write("-- generated by tools/bench/prep_dlx.py -- do not edit\nreturn {\n") fh.write("-- generated by tools/bench/prep_dlx.py -- do not edit\nreturn {\n")
fh.write(f" W={d.W}, H={d.H}, fps={d.fps}, nframes={d.nframes},\n") fh.write(f" W={d.W}, H={d.H}, fps={d.fps}, nframes={NFRAMES},\n")
fh.write(f" k1={d.k1}, k4={d.k4}, dark={dark},\n") fh.write(f" k1={d.k1}, k4={d.k4}, dark={dark},\n")
fh.write(f" cb1_len={cb1.nbytes}, cb4_len={cb4.nbytes}, pal_len={palb.nbytes},\n") fh.write(f" cb1_len={cb1.nbytes}, cb4_len={cb4.nbytes}, pal_len={palb.nbytes},\n")
fh.write(f" stream_len={len(stream)},\n") fh.write(f" stream_len={len(stream)},\n")
@@ -119,5 +148,13 @@ print(f" cb1 {cb1.nbytes} B + cb4 {cb4.nbytes} B expanded, palette {palb.nbytes
print(f" non-SKIP blocks/frame: median {np.median(ns):.1f}% " print(f" non-SKIP blocks/frame: median {np.median(ns):.1f}% "
f"p90 {np.percentile(ns,90):.1f}% max {ns.max():.1f}%") f"p90 {np.percentile(ns,90):.1f}% max {ns.max():.1f}%")
print(f" darkest palette entry: index {dark} -> {tuple(render(I)[dark])}") print(f" darkest palette entry: index {dark} -> {tuple(render(I)[dark])}")
print(f" 4-byte record alignment cost {pad} B over {d.nframes} frames " # A DLX2 container already carries this padding (FINDINGS 28.3 closed, session
f"({pad / d.nframes:.2f} B/frame = {pad / d.nframes * d.fps:.0f} B/s)") # 9), so the realignment above re-derives bytes that were already there and the
# loader is doing no work. On a DLX1 container it is load-bearing: 94 of 120
# record starts land on odd addresses, and each one is an address error.
src_bad = sum(1 for (o, _) in d.frames[:NFRAMES] if (o - 4) % 4)
print(f" 4-byte record alignment cost {pad} B over {NFRAMES} frames "
f"({pad / NFRAMES:.2f} B/frame = {pad / NFRAMES * d.fps:.0f} B/s)")
print(f" source container is DLX{d.version}: {src_bad}/{NFRAMES} record starts "
f"unaligned" + (" -- this loader is what makes it decodable"
if src_bad else " -- the container carries its own padding"))
+12 -3
View File
@@ -22,11 +22,20 @@ from dlx import DLX
ap = argparse.ArgumentParser() ap = argparse.ArgumentParser()
ap.add_argument("container") ap.add_argument("container")
ap.add_argument("--snap", default="tmp/snap_decode") ap.add_argument("--snap", default="tmp/snap_decode")
# The harness can only load as much of a container as fits in the emulated
# machine's RAM, so it may have decoded a PREFIX (tools/bench/prep_dlx.py
# --ram). Compare against the same prefix, or the reference runs ahead of the
# 68000 and reports a mismatch that is an artefact of the rig.
ap.add_argument("--nframes", type=int, default=None,
help="frames the 68000 actually decoded (default: all)")
a = ap.parse_args() a = ap.parse_args()
d = DLX(a.container) d = DLX(a.container)
NF = a.nframes if a.nframes is not None else d.nframes
if NF > d.nframes:
sys.exit(f"--nframes {NF} exceeds the container's {d.nframes}")
canvas = np.zeros((d.H, d.W), np.uint8) canvas = np.zeros((d.H, d.W), np.uint8)
for f in range(d.nframes): for f in range(NF):
d.paint(canvas, f) d.paint(canvas, f)
pal = d.pal.astype(int) pal = d.pal.astype(int)
@@ -52,7 +61,7 @@ else:
bad = diff.any(2) bad = diff.any(2)
by, bx = np.where(bad) by, bx = np.where(bad)
blocks = sorted(set(zip((by//4).tolist(), (bx//4).tolist()))) blocks = sorted(set(zip((by//4).tolist(), (bx//4).tolist())))
fail.append(f"3. frame {d.nframes-1} not pixel-exact: {bad.sum()} px in " fail.append(f"3. frame {NF-1} not pixel-exact: {bad.sum()} px in "
f"{len(blocks)} blocks differ, maxdiff {diff.max()}; " f"{len(blocks)} blocks differ, maxdiff {diff.max()}; "
f"first block (by={blocks[0][0]}, bx={blocks[0][1]})") f"first block (by={blocks[0][0]}, bx={blocks[0][1]})")
@@ -60,7 +69,7 @@ for x in fail:
print("FAIL " + x) print("FAIL " + x)
if fail: if fail:
sys.exit(1) sys.exit(1)
print(f"OK {d.nframes} frames decoded on the 68000, final frame pixel-exact " print(f"OK {NF} frames decoded on the 68000, final frame pixel-exact "
f"against tools/encoder/dlx.py") f"against tools/encoder/dlx.py")
print(f" {d.W}x{d.H}, {d.nb} blocks/frame, k1={d.k1} k4={d.k4}, " print(f" {d.W}x{d.H}, {d.nb} blocks/frame, k1={d.k1} k4={d.k4}, "
f"all four block modes exercised") f"all four block modes exercised")
+19 -4
View File
@@ -23,8 +23,15 @@ class DLX:
def __init__(self, path): def __init__(self, path):
self.raw = open(path, "rb").read() self.raw = open(path, "rb").read()
b = self.raw b = self.raw
if b[:4] != b"DLX1": # DLX2 pads every frame record up to a 4-byte boundary; DLX1 lays them
raise ValueError(f"{path}: not a DLX1 container") # end to end. On a 68000 that is not a slow read but an ADDRESS ERROR
# (FINDINGS 28.3), so the padding is part of the format, not a loader
# convenience -- but DLX1 containers stay readable, because every
# measurement in FINDINGS 28-31 was taken on one.
if b[:4] not in (b"DLX1", b"DLX2"):
raise ValueError(f"{path}: not a DLX container")
self.version = int(b[3:4])
self.aligned = self.version >= 2
(self.W, self.H, self.fps, self.nframes, (self.W, self.H, self.fps, self.nframes,
self.k1, self.k4) = struct.unpack(">HHHHHH", b[4:16]) self.k1, self.k4) = struct.unpack(">HHHHHH", b[4:16])
off_pal, off_cb1, off_cb4, off_frm = struct.unpack(">IIII", b[16:32]) off_pal, off_cb1, off_cb4, off_frm = struct.unpack(">IIII", b[16:32])
@@ -41,14 +48,22 @@ class DLX:
self.mode_bytes = (self.nb * 2 + 7) // 8 self.mode_bytes = (self.nb * 2 + 7) // 8
# frame directory: (offset of the mode header, payload length) # frame directory: (offset of the mode header, payload length)
if self.aligned and off_frm % 4:
raise ValueError(f"{path}: DLX2 frame stream starts at {off_frm}, "
f"which is not 4-byte aligned")
self.frames = [] self.frames = []
p = off_frm p = off_frm
for _ in range(self.nframes): for _ in range(self.nframes):
(n,) = struct.unpack(">I", b[p:p + 4]) (n,) = struct.unpack(">I", b[p:p + 4])
self.frames.append((p + 4, n)) self.frames.append((p + 4, n))
p += 4 + n p += 4 + n
if p != len(b): if self.aligned:
raise ValueError(f"{path}: {len(b) - p} trailing bytes after " p += -p % 4 # skip the pad to the next record
# The writer does not pad after the LAST record -- nothing follows it --
# so `p` may have advanced past the end by up to 3 bytes there.
slack = len(b) - p
if not (slack == 0 or (self.aligned and -3 <= slack < 0)):
raise ValueError(f"{path}: {slack} trailing bytes after "
f"{self.nframes} frames") f"{self.nframes} frames")
def modes(self, f): def modes(self, f):
+24 -7
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Encode one scene to the DLX bitstream, at a chosen quality profile. """Encode one scene to the DLX bitstream, at a chosen quality profile.
python3 tools/encoder/encode.py <frames_dir> <out.dlx> [--profile sasi|scsi] python3 tools/encoder/encode.py <frames_dir> <out.dlx> [--profile scsi]
[--lam N] [--fps 12] [--preview out.png] [--lam N] [--fps 12] [--preview out.png]
[--fixed-lam] [--rc-floor profile|open] [--fixed-lam] [--rc-floor profile|open]
@@ -16,7 +16,7 @@ Container (little-endian is WRONG here -- the 68000 is big-endian, so every
multi-byte field is big-endian and the decoder can read it with a plain move.w): multi-byte field is big-endian and the decoder can read it with a plain move.w):
header, 32 bytes header, 32 bytes
0 'DLX1' magic 0 'DLX2' magic ('DLX1' = the same, unaligned; still read)
4 u16 width, u16 height 4 u16 width, u16 height
8 u16 fps, u16 nframes 8 u16 fps, u16 nframes
12 u16 k1, u16 k4 codebook sizes 12 u16 k1, u16 k4 codebook sizes
@@ -25,7 +25,9 @@ multi-byte field is big-endian and the decoder can read it with a plain move.w):
20 u32 cb1 offset (k1 * 16 bytes of palette indices) 20 u32 cb1 offset (k1 * 16 bytes of palette indices)
24 u32 cb4 offset (k4 * 4 bytes) 24 u32 cb4 offset (k4 * 4 bytes)
28 u32 frames offset 28 u32 frames offset
then, per frame: then, per frame, each record starting on a 4-BYTE BOUNDARY (0-3 zero pad
bytes before it; a 68000 takes an address error, not a slow read, on an odd
`move.l` -- FINDINGS 28.3):
u32 payload length, then u32 payload length, then
ceil(nblocks*2/8) bytes of 2-bit mode headers, MSB-first, block raster order ceil(nblocks*2/8) bytes of 2-bit mode headers, MSB-first, block raster order
then payloads in block order: V1 -> 1 byte, V4 -> 4 bytes, RAW -> 16 bytes then payloads in block order: V1 -> 1 byte, V4 -> 4 bytes, RAW -> 16 bytes
@@ -83,7 +85,7 @@ def main():
global _IDX_BYTES global _IDX_BYTES
ap = argparse.ArgumentParser() ap = argparse.ArgumentParser()
ap.add_argument("frames_dir"); ap.add_argument("out") ap.add_argument("frames_dir"); ap.add_argument("out")
ap.add_argument("--profile", choices=list(RC.PROFILES), default="sasi") ap.add_argument("--profile", choices=list(RC.PROFILES), default="scsi")
ap.add_argument("--lam", type=float, default=None) ap.add_argument("--lam", type=float, default=None)
ap.add_argument("--fps", type=int, default=12) ap.add_argument("--fps", type=int, default=12)
ap.add_argument("--iters", type=int, default=16) ap.add_argument("--iters", type=int, default=16)
@@ -165,19 +167,34 @@ def main():
off_cb1 = off_pal + len(pal_b) off_cb1 = off_pal + len(pal_b)
off_cb4 = off_cb1 + len(cb1_b) off_cb4 = off_cb1 + len(cb1_b)
off_frm = off_cb4 + len(cb4_b) off_frm = off_cb4 + len(cb4_b)
hdr = (b"DLX1" + struct.pack(">HHHHHH", W_, H_, a.fps, len(idx), k1, k4) # DLX2: every frame record starts on a 4-byte boundary, including the
# first. Payload lengths are arbitrary, so end-to-end records land on odd
# addresses -- and `move.l (a0)+` at an odd address is an ADDRESS ERROR on
# a 68000, not a slow read. It vectors into the IPL and looks exactly like
# an infinite loop (FINDINGS 28.3). tools/bench/prep_dlx.py has been
# realigning at load time; the container now carries it.
tbl_pad = -off_frm % 4
off_frm += tbl_pad
hdr = (b"DLX2" + struct.pack(">HHHHHH", W_, H_, a.fps, len(idx), k1, k4)
+ struct.pack(">IIII", off_pal, off_cb1, off_cb4, off_frm)) + struct.pack(">IIII", off_pal, off_cb1, off_cb4, off_frm))
assert len(hdr) == 32, len(hdr) assert len(hdr) == 32, len(hdr)
frm_pad = 0
with open(a.out, "wb") as fh: with open(a.out, "wb") as fh:
fh.write(hdr); fh.write(pal_b); fh.write(cb1_b); fh.write(cb4_b) fh.write(hdr); fh.write(pal_b); fh.write(cb1_b); fh.write(cb4_b)
for p in frames: fh.write(b"\0" * tbl_pad)
for i, p in enumerate(frames):
fh.write(struct.pack(">I", len(p))); fh.write(p) fh.write(struct.pack(">I", len(p))); fh.write(p)
if i + 1 < len(frames): # nothing follows the last record
n = -(4 + len(p)) % 4
fh.write(b"\0" * n); frm_pad += n
total = os.path.getsize(a.out) total = os.path.getsize(a.out)
vid = sum(len(p) + 4 for p in frames) vid = sum(len(p) + 4 for p in frames) + frm_pad
print(f" wrote {a.out}: {total} B " print(f" wrote {a.out}: {total} B "
f"(header+tables {total-vid} B, video {vid} B)") f"(header+tables {total-vid} B, video {vid} B)")
print(f" DLX2 4-byte record alignment: {frm_pad} B over {len(frames)} frames "
f"({frm_pad/len(frames):.2f} B/frame = {frm_pad/len(frames)*a.fps:.0f} B/s)")
print(f" {vid/len(idx):.0f} B/frame -> {vid/len(idx)*a.fps/1024:.1f} KB/s video" print(f" {vid/len(idx):.0f} B/frame -> {vid/len(idx)*a.fps/1024:.1f} KB/s video"
f" + {RC.AUDIO_KBPS} KB/s audio = {vid/len(idx)*a.fps/1024+RC.AUDIO_KBPS:.1f} KB/s") f" + {RC.AUDIO_KBPS} KB/s audio = {vid/len(idx)*a.fps/1024+RC.AUDIO_KBPS:.1f} KB/s")
print(f" PSNR {r['psnr']:.2f} dB palette ceiling {r['pal']:.2f} dB " print(f" PSNR {r['psnr']:.2f} dB palette ceiling {r['pal']:.2f} dB "
+17 -10
View File
@@ -49,18 +49,25 @@ import vq_hybrid as H
# ~216% of the frame budget on a 68000 and even LZ4 is ~54%. See FINDINGS 17. # ~216% of the frame budget on a 68000 and even LZ4 is ~54%. See FINDINGS 17.
# The rates below are therefore RAW payload, no entropy coding. # The rates below are therefore RAW payload, no entropy coding.
# #
# The two profiles are the SAME codec, decoder and bitstream -- only `lam` differs.
# `lam` here is a FLOOR, not a setting: encode.py rate-controls by default and # `lam` here is a FLOOR, not a setting: encode.py rate-controls by default and
# bisects lam per frame in [lam, LAM_CLIFF] to keep under `kbps`. The floor is # bisects lam per frame in [lam, LAM_CLIFF] to keep under `kbps`. The floor is
# what a quiet frame is allowed to spend, so rate control can only ever spend # what a quiet frame is allowed to spend, so rate control can only ever spend
# less than session 5's fixed-lam encoder did. FINDINGS 27. # less than session 5's fixed-lam encoder did. FINDINGS 27.
#
# THE `sasi` PROFILE IS GONE (session 9, USER DECISION). It was dropped on
# CAPACITY, not bandwidth: a SASI volume on this machine tops out at 40 MB, and
# the 22.8 minutes of unique scene footage on the source Blu-ray is 147 MB even
# at the 110 KB/s the profile targeted -- more than the whole 4-unit SASI
# address space, with nothing left for Human68k or the game. FINDINGS 32.
#
# That leaves ONE profile, which is also the end of the two-quality-mode
# decision of session 2. The 110 KB/s RATE POINT may still return under another
# name: a 1x SCSI CD-ROM sustains ~150 KB/s, below this profile, and CD-ROM is
# the only period medium with the capacity for the span-heavy stream. That is
# deferred to the blocked disk benchmark and the DMA-vs-PIO check (docs/
# BENCHMARK.md, FINDINGS 29.5), because every bandwidth figure here is folklore
# until one of them lands.
PROFILES = { PROFILES = {
"sasi": dict(kbps=110, lam=60.0, k1=256, k4=256,
desc="stock 10MHz ACE/EXPERT, SASI",
quality="36.9 dB on 00020 / 29.6 dB on 00146 / 27.2 dB on the "
"Singe window at 109.5 KB/s (session 5's fixed lam "
"gave 27.8 dB there, but at 137.4 KB/s)",
util="~105 KB/s = 35% of the pessimistic 300 KB/s SASI figure"),
"scsi": dict(kbps=280, lam=10.0, k1=256, k4=256, "scsi": dict(kbps=280, lam=10.0, k1=256, k4=256,
desc="Super/XVI, or CZ-6BS1 board in a 10MHz machine", desc="Super/XVI, or CZ-6BS1 board in a 10MHz machine",
quality="39.4 dB on 00020 / 32.3 dB on 00146 / 29.9 dB on the " quality="39.4 dB on 00020 / 32.3 dB on 00146 / 29.9 dB on the "
@@ -82,9 +89,9 @@ PROFILES = {
LAM_CLIFF = 800.0 LAM_CLIFF = 800.0
# Ceiling on the CYCLE search. mu prices a cycle in the same units lam prices a # Ceiling on the CYCLE search. mu prices a cycle in the same units lam prices a
# byte, so the scale that matters is set by their ratio: at the `sasi` floor of # byte, so the scale that matters is set by their ratio: at a lam floor of 60
# lam=60, mu=0.2 makes a V1 block's 300 cycles cost what its 1 payload byte # (the retired `sasi` profile's, and the highest this codec has shipped),
# costs. MU_CLIFF=100 is three decades past that: a V1 block priced at 30,000 # mu=0.2 makes a V1 block's 300 cycles cost what its 1 payload byte costs. MU_CLIFF=100 is three decades past that: a V1 block priced at 30,000
# distortion units. # distortion units.
# #
# It does NOT freeze the picture, and that is the point. At MU_CLIFF a block # It does NOT freeze the picture, and that is the point. At MU_CLIFF a block