diff --git a/README.md b/README.md index d2ed51f..952de48 100644 --- a/README.md +++ b/README.md @@ -37,12 +37,17 @@ tools/analysis/ measurement scripts, numbered in the order they were written demonstrates that the two-display-path plan of FINDINGS 24.5/25.6 corrupts 70 of 120 frames (FINDINGS 28.1). 11 scores a container against the MEASURED per-mode block - costs without needing MAME. + costs without needing MAME; 12 prices the literal-span mode of + FINDINGS 30 against those same mode maps, and prints whether a + scene cut still fits at 12fps. tools/bench/ MAME Lua injection harness + 68000 benchmark sources. `check.sh` re-runs both display regression tests (~40 s). `blit.s`/`blit.lua` time the full-frame GVRAM blit on the 68000 itself (FINDINGS 24) — not part of check.sh, because wall timings would make the green-light check host-sensitive. + `span.sh` (prep_spans.py + span.lua + blit.s v5/v6) measures + the literal-span mode the same way (FINDINGS 30, ~25 s); it + also asserts all 23 timing configs drew a pixel-exact frame. `crtc_mode.lua` is the single source of truth for CRTC R00-R08 and R20 — do not write CRTC values anywhere else. `prep_dlx.py`/`decode.lua`/`verify_decode.py` load, time and diff --git a/docs/FINDINGS.md b/docs/FINDINGS.md index e902dc2..7c713ca 100644 --- a/docs/FINDINGS.md +++ b/docs/FINDINGS.md @@ -1428,10 +1428,17 @@ without closing it. ## 29. Trading bytes for cycles: the bus has 4x the headroom the CPU has (session 7) -> **STATUS: DERIVED, NOT MEASURED.** No 68000 has executed a span decoder. The -> per-pixel figure it rests on *is* measured (FINDINGS 24 V1) but at full row -> width; the per-span overhead is hand-derived. Treat every number below as a -> hypothesis with a test attached, not as a result. FINDINGS 4 is why. +> **SUPERSEDED IN PART BY 30, which measured it.** The mode survives and the +> conclusion holds, but every number in this section moved: a span costs 43.7 +> cycles + 9.152/pixel *only* in an encoder-assisted format (the obvious +> decoder is 97.9 + 10.46), spans beat V1 from runs of 4 blocks and not 2, and +> the re-priced trade-off is 52.0% median / 10 misses, not 43.0% / 8. Read 30's +> tables over 29.3's. 29.5's other three items are still open, and 29.6 stands. +> +> **STATUS AT THE TIME: DERIVED, NOT MEASURED.** No 68000 had executed a span +> decoder. The per-pixel figure it rests on *is* measured (FINDINGS 24 V1) but +> at full row width; the per-span overhead was hand-derived. FINDINGS 4 is why +> it was labelled and then tested rather than believed. FINDINGS 28 leaves the project CPU-bound while the **bus sits 4x idle**: `sasi` spends 110 KB/s of a 488 KB/s pipe. That asymmetry is exploitable, because the @@ -1534,3 +1541,140 @@ It cannot be settled in MAME: like the SCSI/SASI devices (BENCHMARK.md), the HD63450 is a functional model, so a timing number out of it would measure the emulator's scheduler. It needs hand-derivation against the datasheet plus real hardware — the same three-tier approach the disk benchmark already documents. + +## 30. The span, measured: the mode survives, and it is an encoder format (session 8) + +FINDINGS 29 priced a new decoder mode at `4 * (50 + 4L*9.08)` cycles and marked +the whole section DERIVED. This is the measurement. `tools/bench/blit.s` gained +two span variants, `tools/bench/prep_spans.py` generates one stream per span +length, `tools/bench/span.lua` times them, and `tools/bench/span.sh` runs the +lot, and the whole thing takes about 25 seconds. + +Same scope as every 68000 figure since FINDINGS 24: instruction cycles against +MAME's zero-wait-state GVRAM, interrupts masked. A **lower bound**, not a +prediction. + +### 30.1 What was measured +Twelve `v5` configs and eleven `v6` configs, each cutting the **same** 256x192 +frame into spans of a different length, so the work differs only in how finely +it is cut. Regressing `cycles = A*spans + B*pixels` over a set reads the +per-span overhead and the per-pixel cost straight off. + +Every config draws the whole picture, the picture is cleared before each run and +snapshotted after, and all 23 snapshots are checked pixel-exact by +`tools/bench/verify_frame256.py`. A config cannot time fast by writing nothing. + +| | per span | per pixel | fit error | +|---|---:|---:|---:| +| **v5** — decoder handed `(x, npix)`, works the copy out | **97.9** | **10.459** | ±1.4%, and only on spans that are a whole number of bursts | +| **v6** — encoder hands it an address and a jump | **43.7** | **9.152** | **±0.3% over all 11 lengths** | +| *29's assumption* | *50.0* | *9.080* | — | + +**29's arithmetic was right about a format nobody had written yet.** v6 hits it +almost exactly; v5 — the obvious decoder, and the one 29 was describing — is +2.24x dearer per span and 14% dearer per pixel. + +### 30.2 Why the difference is a format difference, not an optimisation +v5's record is `(x, npix)`, so the decoder computes the destination, divides +`npix` into 16-pixel bursts, and handles the 0..15 remainder: about 122 cycles +of arithmetic and branching per span before a single pixel moves. All of it is +known at encode time. + +v6's record is `{u32 absolute GVRAM address, u16 jump displacement}` and nothing +else. The displacement jumps into an unrolled chain of eleven 24-pixel copy +units, so a span of any supported length is straight-line code with no loop, no +remainder, and no address arithmetic — `move.l (a0)+,a2` / `move.w (a0)+,d0` / +`jmp v6ch(pc,d0.w)`, then `movem.l` pairs. GVRAM is at $C00000 on every X68000, +so absolute destinations are a legitimate thing to bake into a stream. + +Two consequences of that format, both cheap: +- **Span lengths are multiples of 24 pixels** and a run pads up to it. The + padding costs bytes and its own pixels, nothing else, and it is *correct on + screen*: a literal span carries true pixels of the current frame, so painting + a clean neighbour is a no-op visually. +- **A span may overrun the visible 256 pixels of its row by up to 23.** Free: + the line stride is 1024 bytes and only the first 512 are displayed, so the + overrun lands in the invisible half of the line. + +### 30.3 The remainder path is where a short span actually dies +v5's cost per span, measured, against its length: + +| span | 4 px | 8 px | 12 px | 16 px | 20 px | 24 px | 32 px | +|---|---:|---:|---:|---:|---:|---:|---:| +| cycles/span | 180.3 | 240.9 | 296.3 | 261.8 | 347.7 | 401.9 | 430.7 | +| cycles/pixel | 45.08 | 30.11 | 25.46 | **16.36** | 17.65 | 17.27 | **13.46** | + +A 12-pixel span costs *more* than a 16-pixel one. Everything below the 16-pixel +burst width goes through `move.l`/`move.w` at roughly 10 cycles a pixel plus the +per-span overhead, and 29's warning that "short spans are flattered" was +correct — but the fix is to pad them up to a burst, not to avoid them. v6 has no +remainder path at all, which is most of why its fit is linear to 0.3%. + +### 30.4 Registers are the reason the per-pixel cost moved +FINDINGS 24's 9.08 cycles/pixel came from a fixed blit with 12 registers free +for `movem.l` and no live state. A span decoder keeps a stream pointer, a +destination and counters live, so v5 can spare only 8 registers per burst — 32 +bytes instead of 48 — and pays 10.46 cycles/pixel for it. v6 gets back to 12 +registers precisely because the encoder holds the state instead, and lands at +9.152. **The per-pixel figure is a function of how much the decoder has to +remember**, which is not something the FINDINGS 24 measurement could have shown. + +Two smaller results, both cheap and both worth having on the record: +- **Odd-`x` alignment is free.** Spans starting at an odd pixel run their bursts + at `addr mod 4 == 2` and cost 259.0 cycles/span against 261.8 aligned — inside + the timing granularity. The 68000's 16-bit bus does not care, as expected; + now it is measured rather than assumed. +- **A full-row span is 154 cycles per 4x4 block**, the floor this mode can + reach, against V1's measured 299.9. + +### 30.5 Re-pricing: the trade holds, and it is smaller +`tools/analysis/12_span_tradeoff.py` now runs on measured constants. Same greedy +(buy the best cycles-saved-per-byte until the bus budget is gone), same +unmodified mode maps, same Singe window: + +| | today | 29 (derived) | **30 (measured)** | +|---|---:|---:|---:| +| `sasi` median frame | 74.4% | 43.0% | **52.0%** | +| `sasi` worst frame | 136.2% | 106.2% | **108.7%** | +| `sasi` frames missing | 37/120 | 8/120 | **10/120** | +| `sasi` bitrate | 101.7 KB/s | 453.2 | **448.0 KB/s** | +| `scsi` median frame | 94.9% | 69.4% | **74.6%** | +| `scsi` frames missing | 51/120 | 18/120 | **25/120** | + +And the break-even moved. Cycles per 4x4 block in a run of L blocks, v6, with +each of the run's 4 spans padded to a whole 24-pixel unit: + +| L | 1 | 2 | 4 | 8 | 16 | 64 | +|---|---:|---:|---:|---:|---:|---:| +| cycles/block | 1053 | 527 | **263** | 242 | 176 | 154 | + +So a run beats all-V1 (299.9) **from L=4 up**, not from L=2 as 29.3 claimed, and +runs of 1-3 blocks all cost the same 1053 cycles because they pad to the same +single unit. A cost-aware mode decision should not offer a span below 4 blocks +at all. + +### 30.6 29.4 survives: a scene cut still fits at 12fps +Mixing a fraction `x` of a 100%-changed frame as full-row spans against V1 for +the rest, on measured costs (154 cycles and 33.4 bytes per block): + +- CPU needs `x >= 0.196` +- the 40,977 B/frame bus budget allows `x <= 0.373` + +The interval is not empty — narrower than 29.4's 0.19..0.39, same conclusion. +FINDINGS 28.5's "a scene cut cannot fit" was a ceiling of the bitstream, not of +the machine, and that now rests on a measurement. `12_span_tradeoff.py` prints +this arithmetic and will say so if it ever stops being true. + +### 30.7 What this does NOT settle +The three remaining items of 29.5 are unchanged and are now **more** load-bearing, +because the measured design runs at 448 KB/s of a 488 KB/s pipe rather than 453: +re-run the ring-buffer simulation at that rate, confirm the 4 Mbps figure's +provenance, and confirm DMA rather than PIO. A PIO fallback would put a 448 KB/s +transfer back on the CPU this mode exists to relieve. + +Also unmeasured: **the parse cost of a span-heavy stream**. Every figure here +times the copy. The 68000 also has to read the mode map and dispatch: the +re-priced `sasi` stream buys 8773 spans across 120 frames, a mean of 73 a frame, +and each one's three-instruction dispatch is inside the fitted 43.7 — but the +mode-map walk that decides a span exists is not. `decode.s` does not implement +spans yet. diff --git a/docs/STATUS.md b/docs/STATUS.md index 10527d7..2993e5d 100644 --- a/docs/STATUS.md +++ b/docs/STATUS.md @@ -1,60 +1,48 @@ -# Status & next-session handoff — end of session 7 (2026-08-23) +# Status & next-session handoff — session 8 (2026-08-23) -## NEXT SESSION: measure a span, then make the mode decision cost-aware +## Where this stands -The decoder exists, it is pixel-exact, and **it does not fit**. On the worst -sustained window at `sasi` it costs a mean of **81.7% of a 12fps frame** and -**31% of frames exceed 100%** (`scsi`: 94.9% median, 42% miss). FINDINGS 28. -CPU is the binding constraint now — the first time in this project. +The decoder exists, it is pixel-exact, and **it does not fit**: mean 81.7% of a +12fps frame on the worst sustained window at `sasi`, 31% of frames over budget +(`scsi`: 94.9% median, 42% miss). FINDINGS 28. CPU is the binding constraint. -**Two levers, and the cheap one has to be measured first.** +Session 7 proposed two levers and session 8 measured the cheap one first. -*Lever A — spend bandwidth to buy cycles.* The bus sits 4x idle: `sasi` uses 110 -KB/s of 488. Every codec decision was made when bytes were scarce, so each one -trades cycles to save them, and the cheapest thing a 68000 can be handed is the -most expensive thing to store — **word-expanded pixels in row-linear runs**. -Adding one mode, a per-row span of literal words `movem.l`-ed straight from the -stream buffer into GVRAM, prices out at (FINDINGS 29, `12_span_tradeoff.py`): +**Lever A — spend bandwidth to buy cycles — is real, and it is an encoder +format.** A row-linear span of word-expanded literals measures **43.7 cycles per +span + 9.152 per pixel** (FINDINGS 30, `tools/bench/span.sh`), which is what +FINDINGS 29 assumed — but only when the *encoder* hands the decoder an absolute +GVRAM address and a jump displacement into an unrolled copy chain. The obvious +decoder, handed `(x, npix)` and left to work the copy out, is 97.9 + 10.46 and +2.2x dearer on a short span. Re-priced against the unchanged mode maps: -| | today | + literal spans | -|---|---:|---:| -| median frame | 74.4% | **43.0%** | -| worst frame | 136.2% | **106.2%** | -| frames missing | **37/120** | **8/120** | -| bitrate | 101.7 KB/s | 453.2 KB/s (bus 488) | +| | today | 29 (derived) | **30 (measured)** | +|---|---:|---:|---:| +| `sasi` median frame | 74.4% | 43.0% | **52.0%** | +| `sasi` worst frame | 136.2% | 106.2% | **108.7%** | +| `sasi` frames missing | 37/120 | 8/120 | **10/120** | +| bitrate | 101.7 KB/s | 453.2 | **448.0 KB/s** (bus 488) | -**This is DERIVED, not measured, and it is load-bearing — so measure it first.** -Extend `tools/bench/blit.s` with a span variant and time it against run length. -The 9.08 cycles/pixel it rests on is real (FINDINGS 24 V1) but was measured at -full row width with 12-register bursts; short and oddly-aligned spans cannot -burst as well and are flattered by the model. If spans come in near the derived -figure, the whole mode set changes and lever B optimises over different modes — -which is exactly why this goes first. FINDINGS 29.5 lists the other three things -that have to hold, of which **confirming DMA vs PIO is the cheapest and now the -most consequential**: at 453 KB/s a PIO fallback puts the transfer back on the -CPU this is trying to relieve. +Break-even moved with it: a run beats all-V1 **from 4 blocks up**, not 2. And +29.4 survives — a scene cut needs `x >= 0.196` of the frame as spans and the bus +allows `x <= 0.373`, so it fits at 12fps. -*Lever B — stop buying modes the CPU cannot afford.* `vq_hybrid.decide()` -minimises `D + lam*R` — distortion against BYTES — on a machine whose binding -budget is CYCLES, and the two are not proportional: +**Lever B — stop buying modes the CPU cannot afford — is untouched.** +`vq_hybrid.decide()` still minimises `D + lam*R`, distortion against BYTES, on a +machine whose binding budget is CYCLES: -| mode | payload bytes | measured cycles | cycles per byte | +| mode | payload bytes | cycles | cycles per byte | |---|---:|---:|---:| | SKIP | 0 | 13 (clustered) | — | | V1 | 1 | 300 | 300 | | V4 | 4 | 448 | 112 | | RAW | 16 | 400 | 25 | -| *word-expanded literal block* | *32* | *~240 (derived)* | *7.5* | +| **span, per 4x4 block in a run of L** | **32** | **1053/L, floor 154** | **~5** | -V4 is **25% of blocks and 50% of the cycles**. The lagrangian charges it 4x a V1 -block; the CPU charges it 1.49x. Note the last row: a literal block is cheaper -than **every** codebook mode, and pixel-exact — the codebook is a byte -optimisation that now costs cycles (FINDINGS 29.2). +V4 is 25% of blocks and 50% of the cycles. The lagrangian charges it 4x a V1 +block; the CPU charges it 1.49x. -**The work, in order:** - -0. **Measure the span cost on the 68000** (lever A above). Cheap, and everything - below optimises over whatever mode set it leaves. +## The work, in order 1. **Add a cycle term to the mode decision.** `decide()` already builds a cost matrix of `error + lam * bytes` per mode per block; add `+ mu * cycles`, @@ -86,10 +74,11 @@ optimisation that now costs cycles (FINDINGS 29.2). `tools/bench/decode.lua`. 3b. **Know which misses are yours to fix before starting.** Re-coding every - non-SKIP block as V1 is the floor any mode assignment can reach, and it - still misses 11 frames at `sasi` and 12 at `scsi` — every frame above ~90% - non-SKIP. So the cost-aware decision can reach about three quarters of the - misses (26 of 37 at `sasi`) and the rest are item 4. FINDINGS 28.7. + non-SKIP block as V1 is the floor any mode assignment *of the current mode + set* can reach, and it still misses 11 frames at `sasi` and 12 at `scsi` — + every frame above ~90% non-SKIP. So the cost-aware decision can reach about + three quarters of the misses (26 of 37 at `sasi`) and the rest need item 4. + FINDINGS 28.7. 3c. **Buy RAW, not V4, wherever the bytes allow.** RAW is 400 cycles against V4's 448 *and* is pixel-exact, so on the CPU axis V4 is strictly dominated — @@ -97,15 +86,18 @@ optimisation that now costs cycles (FINDINGS 29.2). `sasi` cannot afford it, so expect the cycle ceiling to cost `sasi` more quality even though it costs `sasi` fewer cycles. FINDINGS 28.8. -4. **Scene cuts: 28.5 said impossible, 29.4 reopened it.** An all-V1 frame — - the cheapest full redraw the *current* mode set allows — is 110.5% of budget, - so no mode assignment fits a 100%-changed frame. With literal spans the - arithmetic changes: CPU needs at least 19% of the frame sent as spans, the - bus allows up to 39%, **and that interval is not empty**. So 28.5 was a - ceiling of the bitstream, not of the machine — *if* lever A measures out. - If it does not, this is still a design decision that needs the user: one late - frame at each cut (the outgoing content is unrelated, so it may be - invisible), a cut spread over two frame times, or 10fps. +4. **Put spans in the bitstream** — the mode is measured and nothing implements + it. This is a container change (`encode.py`, `dlx.py`, `decode.s`), a mode + decision that can see runs rather than blocks, and the 24-pixel quantisation + and row-overrun rules of FINDINGS 30.2. It subsumes item 4 of session 7's + plan: with spans, a scene cut fits. + +5. **The three things 30.7 leaves open, now more load-bearing than before**, + because the span design runs at 448 KB/s of a 488 KB/s pipe: re-run the + ring-buffer simulation at that rate (FINDINGS 21 was established at 110 and + 280), confirm the provenance of the 4 Mbps figure, and **confirm DMA rather + than PIO** — a PIO fallback puts a 448 KB/s transfer back on the CPU this + whole lever exists to relieve. The DMA check is the cheapest of the three. **Do not start by hand-optimising `decode.s`.** The hand-derived timings agree with the measurements to 0.5% on V1 and 1% on RAW (FINDINGS 28.4), so the @@ -116,6 +108,33 @@ saves ~16 of 448 cycles. --- +## What session 8 settled + +1. **The span is measured: 43.7 cycles/span + 9.152/pixel, fitted to 0.3% over + eleven span lengths.** `tools/bench/blit.s` v5/v6, `prep_spans.py`, + `span.lua`, driven by `tools/bench/span.sh` (~25 s, not in `check.sh` + because it is a wall timing). FINDINGS 30. +2. **Only in an encoder-assisted format.** `{u32 absolute GVRAM address, u16 + jump displacement}` into an unrolled chain, versus `(x, npix)` and a decoder + that works it out: 43.7 + 9.152 against 97.9 + 10.46. All the arithmetic a + span decoder would do per frame is known at encode time. FINDINGS 30.2. +3. **The per-pixel cost is a function of register pressure**, which FINDINGS 24 + could not have shown: 9.08 was a fixed blit with 12 registers free, v5 can + spare 8 and pays 10.46, v6 gets 12 back by making the encoder hold the state. +4. **Short spans die in the remainder path, and the fix is padding.** A 12-pixel + span costs more than a 16-pixel one in v5. v6 has no remainder path: lengths + are multiples of 24 pixels, padding is free of everything but bytes, and an + overrun past the visible 256 lands in the invisible half of the 1024-byte + line stride. FINDINGS 30.3. +5. **Odd-`x` alignment is free** (259.0 vs 261.8 cycles/span) — expected on a + 16-bit bus, now measured rather than assumed. +6. **The trade is smaller than 29 derived but the conclusion holds**, including + 29.4's reopening of the scene cut. All 23 timing configs also drew a + pixel-exact frame, so nothing here was timed against a decoder that skipped + work. FINDINGS 30.5/30.6. + +--- + ## What session 7 settled 1. **68000 code parses a bitstream and draws frames, pixel-exact.** diff --git a/tools/analysis/12_span_tradeoff.py b/tools/analysis/12_span_tradeoff.py index a37d67a..63bf0af 100644 --- a/tools/analysis/12_span_tradeoff.py +++ b/tools/analysis/12_span_tradeoff.py @@ -12,10 +12,16 @@ This prices ONE new mode against the real mode maps: a per-row SPAN of word-expanded literals, `movem.l`-ed straight from the stream buffer into GVRAM. A run of L horizontally adjacent dirty blocks becomes 4 spans of 4L pixels. -DERIVED, NOT MEASURED (FINDINGS 29). The 9.08 cycles/pixel is measured -(FINDINGS 24 V1) but at full row width with 12-register bursts; SPAN_OVERHEAD is -hand-derived. Short spans are therefore flattered. Measure before believing -- -FINDINGS 29.5 item 1. +MEASURED as of session 8 (FINDINGS 30), on the 68000, with the span decoder in +tools/bench/blit.s v6 and the streams in tools/bench/prep_spans.py: +43.7 cycles per span + 9.152 per pixel, fitting eleven span lengths to within +0.3%. That is the ENCODER-ASSISTED format: the record is an absolute GVRAM +address and a jump displacement into an unrolled copy chain, so the decoder does +no arithmetic per span. The obvious decoder -- handed (x, npix) and left to work +the copy out -- measures 97.9 + 10.46 and is 2.2x dearer on a 24-pixel span (v5). +Span length is therefore a multiple of 24 pixels, and a run pads up to it; the +padding is free of cycles beyond its pixels and correct on screen, because a +literal span carries true pixels of the current frame. The mode maps are NOT re-optimised: this only re-codes regions the encoder already chose to redraw, so it is a lower bound on what a cost-aware encoder @@ -29,12 +35,17 @@ from dlx import DLX FRAME_CYC = 833333.0 # 12fps at 10 MHz AUDIO_KBPS = 7.8 -CYC_PX_ROWLIN = 446286 / 49152. # 9.08, FINDINGS 24 V1 (measured) C_V1, C_V4, C_RAW = 299.9, 448.2, 400.4 # FINDINGS 28.2 (measured) C_SKIP_CLUSTERED, C_SKIP_MIXED = 13.25, 45.0 -SPAN_OVERHEAD = 50.0 # per span, DERIVED +SPAN_OVERHEAD = 43.7 # per span, MEASURED, FINDINGS 30 +CYC_PX_ROWLIN = 9.152 # per pixel, MEASURED, FINDINGS 30 +SPAN_UNIT_PX = 24 # 12 registers of movem.l, one chain unit SPAN_BYTES_PX = 2 # word-expanded: 1 pixel = 1 word -SPAN_HDR = 3 # x, count, and a byte of slack +SPAN_HDR = 6 # u32 GVRAM address + u16 jump displacement + + +def span_px(npix): # a span is a whole number of units + return -(-npix // SPAN_UNIT_PX) * SPAN_UNIT_PX ap = argparse.ArgumentParser() ap.add_argument("container", nargs="?", @@ -78,8 +89,9 @@ for f in range(d.nframes): L = j - i cur_c = sum(BLK_C[int(b)] for b in m[by][i:j]) cur_b = sum(BLK_B[int(b)] for b in m[by][i:j]) - span_c = 4 * (SPAN_OVERHEAD + 4 * L * CYC_PX_ROWLIN) - span_b = 4 * (SPAN_HDR + 4 * L * SPAN_BYTES_PX) + sp = span_px(4 * L) # padded to the chain's 24-pixel unit + span_c = 4 * (SPAN_OVERHEAD + sp * CYC_PX_ROWLIN) + span_b = 4 * (SPAN_HDR + sp * SPAN_BYTES_PX) if span_c < cur_c: cand.append((cur_c - span_c, span_b - cur_b, L)) i = j @@ -108,4 +120,31 @@ print(f" {'bitrate':<24}{bb.mean()*a.fps/1024:>10.1f} KB/s" f"{nb.mean()*a.fps/1024:>13.1f} KB/s") print(f"\nspans taken: {ntaken.sum()} of {ncand.sum()} candidate runs " f"({100*ntaken.sum()/max(ncand.sum(),1):.0f}%) -- the rest priced out by the bus") -print("\nDERIVED, NOT MEASURED: see FINDINGS 29.5 before acting on this.") +brk = next(L for L in range(1, 65) + if 4*(SPAN_OVERHEAD + span_px(4*L)*CYC_PX_ROWLIN) < L*C_V1) +print(f"\nspan cost MEASURED (FINDINGS 30): {SPAN_OVERHEAD:.1f}/span + " + f"{CYC_PX_ROWLIN:.3f}/pixel, {SPAN_UNIT_PX}-pixel units.") +print(f"a run of L blocks beats all-V1 from L={brk} blocks up " + f"({4*(SPAN_OVERHEAD + span_px(4*brk)*CYC_PX_ROWLIN)/brk:.0f} vs {C_V1:.0f} " + f"cycles/block); the floor at a full row is " + f"{4*(SPAN_OVERHEAD + span_px(256)*CYC_PX_ROWLIN)/64:.0f}.") +print("The mode maps are NOT re-optimised, so this is a lower bound on a " + "cost-aware encoder.") + +# FINDINGS 28.5 said a scene cut cannot fit at 12fps: the cheapest full redraw +# the codec's mode set allows is all-V1 at 110.5% of budget. 29.4 reopened that +# on derived span costs; this is the same arithmetic on measured ones. Mix a +# fraction x of a 100%-changed frame as full-row spans, V1 for the rest. +NB = d.nb +row_c = 4 * (SPAN_OVERHEAD + span_px(4 * d.nbx) * CYC_PX_ROWLIN) / d.nbx +row_b = 4 * (SPAN_HDR + span_px(4 * d.nbx) * SPAN_BYTES_PX) / d.nbx +x_cpu = (NB * C_V1 - FRAME_CYC) / (NB * (C_V1 - row_c)) +x_bus = (BYTE_BUD - d.mode_bytes - NB * BLK_B[1]) / (NB * (row_b - BLK_B[1])) +print(f"\nscene cut (100% of blocks change), spans at full row width " + f"({row_c:.0f} cyc, {row_b:.1f} B per block):") +print(f" all-V1 costs {100*NB*C_V1/FRAME_CYC:.1f}% of the frame -- FINDINGS 28.5") +print(f" CPU needs x >= {x_cpu:.3f} of the frame as spans; " + f"the bus allows x <= {x_bus:.3f}") +print(" " + ("the interval is NOT empty: a cut fits at 12fps (FINDINGS 29.4 holds)" + if x_cpu <= x_bus else + "the interval IS empty: a cut does not fit (FINDINGS 28.5 stands)")) diff --git a/tools/bench/blit.s b/tools/bench/blit.s index 1b9053d..3655652 100644 --- a/tools/bench/blit.s +++ b/tools/bench/blit.s @@ -31,6 +31,43 @@ ; the block needs only one base pointer. V4 deliberately scrambles the ; picture (it reads a row-linear source in block order); it is a timing ; probe, which is why the correctness snapshot is taken after V1. +; V5 ROW-LINEAR LITERAL SPANS, the mode priced in FINDINGS 29 and never +; measured. Walks a stream of per-row span records +; row: u16 nspans, then nspans * { u16 x, u16 npix, npix*u16 pixels } +; for 192 rows, copying each span's word-expanded pixels straight from +; the stream buffer into GVRAM. Unlike V1-V4 the work per call is set by +; the STREAM, not by the code, so one variant measures every span length: +; tools/bench/prep_spans.py generates a stream per span length and +; tools/bench/span.lua times them and fits cycles = A*spans + B*pixels. +; The point of the measurement is A -- the per-span overhead FINDINGS 29 +; guessed at 50 cycles -- and how much B degrades from V1's 9.08 when a +; span is too short to burst. Every config covers the whole frame, so +; V5 draws the SAME picture V1 does and can be verified, not just timed. +; +; Bursts are 8 registers (d0-d3/a3-a6 = 32 bytes = 16 pixels), not V1's +; 12: a0/a1/a2 and d4-d7 are all live across a span (stream, row base, +; destination, and three counters). The remainder is copied move.l at a +; time with a leading move.w when it is odd, so a 4-pixel span never +; reaches a movem at all -- which is exactly the case FINDINGS 29's +; full-row-width extrapolation flatters. +; +; V6 the SAME spans with the arithmetic moved into the encoder. V5 measures +; a decoder that is handed (x, npix) and has to work out how to copy it; +; most of its per-span cost is that working-out, and an encoder can do it +; once at build time instead of 12 times a second. V6's record is +; { u32 absolute GVRAM address, u16 jump displacement } -- no row +; structure, no counters, no remainder logic -- and the displacement +; jumps into an unrolled chain of 24-pixel copy units, so a span of any +; supported length is straight-line code with no loop at all. +; GVRAM sits at a fixed $C00000 on every X68000, so absolute destinations +; are a legitimate thing for an encoder to bake in. +; +; Two consequences of the format. Span lengths are multiples of 24 +; pixels, and a span may overrun the 256 visible pixels of its row by up +; to 23 -- harmless, because the line stride is 1024 bytes and only the +; first 512 are displayed, so the overrun lands in the invisible half. +; And with row and remainder handling gone, 12 registers are free again +; (d0-d6/a1/a3-a6), which is why the unit is 24 pixels and not V5's 16. ; ; 12 registers per movem burst (d0-d7/a2-a5 = 48 bytes) is the maximum ; available: a0=src, a1=dst, a6=end sentinel. The row counter lives in the @@ -43,10 +80,14 @@ FLAG = $18000 ; 0 idle / 1 running / $FF done VAR = $18004 ; variant selector, written by Lua ITER = $18008 ; iteration count, written by Lua +SPTR = $1800C ; V5 span stream pointer, written by Lua SRCW = $60000 ; word-expanded frame 192*512 = 96KB SRCB = $80000 ; byte-per-pixel frame 192*256 = 48KB DST0 = $C08000 ; GVRAM + 32*1024 (first picture row) DSTE = $C38000 ; GVRAM + 224*1024 (one past last) +ROWS = 192 ; picture rows a V5 stream describes +V6UNIT = 12 ; bytes of code per V6 chain unit +V6MAX = 11 ; chain units = 11*24 = 264 pixels >= one row org $10000 start: @@ -58,6 +99,10 @@ start: beq v2 cmp.l #4,d0 beq v4 + cmp.l #5,d0 + beq v5 + cmp.l #6,d0 + beq v6 bra v3 ; ---------------------------------------------------------------- V1 @@ -152,5 +197,90 @@ v4blk: movem.l (a0)+,d0-d7 ; 32 bytes = one 4x4 block, expanded bne v4 bra done +; ---------------------------------------------------------------- V5 +; a0 stream, a1 row base, a2 span destination, d7 rows, d6 spans, d5 pixels, +; d4 burst/tail counter. Everything else (d0-d3/a3-a6) is burst payload. +v5: move.l SPTR.l,a0 + lea DST0,a1 + move.w #ROWS-1,d7 +v5row: move.w (a0)+,d6 ; spans in this row + subq.w #1,d6 + bmi.s v5eor ; a row may legitimately have none +v5span: move.w (a0)+,d0 ; x, in pixels + add.w d0,d0 ; one pixel = one word + lea 0(a1,d0.w),a2 + move.w (a0)+,d5 ; pixels in this span + move.w d5,d4 + lsr.w #4,d4 ; 16-pixel bursts + beq.s v5tail + subq.w #1,d4 +v5burst: movem.l (a0)+,d0-d3/a3-a6 ; 32 bytes straight out of the stream + movem.l d0-d3/a3-a6,(a2) + lea 32(a2),a2 + dbra d4,v5burst +v5tail: moveq #15,d4 + and.w d5,d4 ; 0..15 pixels left + beq.s v5eos + lsr.w #1,d4 ; C = odd pixel count + bcc.s v5t2 + move.w (a0)+,(a2)+ +v5t2: subq.w #1,d4 + bmi.s v5eos +v5tl: move.l (a0)+,(a2)+ + dbra d4,v5tl +v5eos: dbra d6,v5span +v5eor: lea 1024(a1),a1 + dbra d7,v5row + subq.l #1,ITER.l + bne v5 + bra done + +; ---------------------------------------------------------------- V6 +; a0 stream, a2 destination, d7 spans remaining; everything else is payload. +v6: move.l SPTR.l,a0 + move.w (a0)+,d7 ; total spans in the frame + subq.w #1,d7 +v6span: move.l (a0)+,a2 ; absolute GVRAM destination + move.w (a0)+,d0 ; (V6MAX - units) * V6UNIT, from the encoder + jmp v6ch(pc,d0.w) +v6ch: + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + movem.l (a0)+,d0-d6/a1/a3-a6 + movem.l d0-d6/a1/a3-a6,(a2) + lea 48(a2),a2 + dbra d7,v6span + subq.l #1,ITER.l + bne v6 + bra done + done: move.l #$FF,FLAG.l ; timer stops here halt: bra.s halt diff --git a/tools/bench/prep_spans.py b/tools/bench/prep_spans.py new file mode 100644 index 0000000..cc4e6ca --- /dev/null +++ b/tools/bench/prep_spans.py @@ -0,0 +1,118 @@ +#!/usr/bin/env python3 +"""Generate V5 span streams for tools/bench/span.lua (FINDINGS 29.5 item 1). + +FINDINGS 29 prices a new decoder mode -- a row-linear run of word-expanded +literal pixels, movem.l'd straight from the stream buffer into GVRAM -- at +`4 * (50 + 4L * 9.08)` cycles for a run of L blocks. Both halves of that are +extrapolations: the 50-cycle per-span overhead is hand-derived, and the 9.08 +cycles/pixel was measured (FINDINGS 24 V1) at FULL ROW WIDTH with 12-register +bursts, which a short span cannot match. This script builds the stimulus that +replaces both numbers with measured ones. + +One stream per span length. Every stream covers the SAME 192x256 picture +completely, so all of them draw an identical, verifiable frame and differ only +in how many spans it is cut into -- which is what lets span.lua regress + cycles = A * spans + B * pixels +across the set and read the per-span overhead off directly. + +Two stream formats, both big-endian, both drawing the same frame. + +v5 -- a decoder handed (x, npix) that works out the copy itself: + per row, 192 rows in order: + u16 nspans + nspans * { u16 x, u16 npix, npix * u16 pixel } + +v6 -- the same spans with that arithmetic moved here, where it is free: + u16 nspans (whole frame; there is no row structure) + nspans * { u32 absolute GVRAM address, u16 jump displacement, + units * 48 bytes of pixels } + Span lengths are multiples of 24 pixels (one chain unit) and the last span + in a row may overrun the visible 256 by up to 23 pixels, which is free: the + line stride is 1024 bytes and only the first 512 are displayed. The jump + displacement selects an entry point into the decoder's unrolled copy chain. + +Pixels are word-expanded with the palette index in the low byte; the high byte +is whatever we put there because gvram_w masks it off (x68k_crtc.cpp:501). +""" +import struct, sys +import numpy as np + +SRC = sys.argv[1] if len(sys.argv) > 1 else "tmp/frame256.bin" +OUT = sys.argv[2] if len(sys.argv) > 2 else "tmp/spans.bin" +META = OUT.replace(".bin", "_meta.lua") + +d = open(SRC, "rb").read() +assert d[:4] == b"DLXR", SRC +W, H = struct.unpack(">HH", d[4:8]) +idx = np.frombuffer(d[8+768:8+768+W*H], np.uint8).reshape(H, W) +assert (W, H) == (256, 192), f"{W}x{H}: span bench assumes the 256x192 picture" + +# (span length in pixels, x of the first span). 4 px = one 4x4 block wide, the +# case the whole FINDINGS 29 argument turns on; 256 = one span per row, the +# case closest to the V1 measurement it extrapolates from. 16u starts at an +# odd x so its bursts run at addr mod 4 == 2: a claim about the 68000's 16-bit +# bus that costs nothing to test and would be embarrassing to assume. +CONFIGS = [(4, 0), (8, 0), (12, 0), (16, 0), (16, 1), (20, 0), (24, 0), + (32, 0), (48, 0), (64, 0), (128, 0), (256, 0)] + +# v6 geometry, and it must match blit.s: 12 registers per movem = 48 bytes = +# 24 pixels per chain unit, 11 units in the chain. +UNITPX, UNITSZ, UNITS = 24, 12, 11 +GVRAM, YOFF, STRIDE = 0xC00000, 32, 1024 + +blob, metas = bytearray(), [] +for P, x0 in CONFIGS: + off = len(blob) + nspans = npix = 0 + for y in range(H): + cuts = [] + x = 0 + if x0: # a short leading span to shift the phase + cuts.append((0, x0)); x = x0 + while x < W: + n = min(P, W - x) + cuts.append((x, n)); x += n + blob += struct.pack(">H", len(cuts)) + for x, n in cuts: + blob += struct.pack(">HH", x, n) + blob += idx[y, x:x+n].astype(">u2").tobytes() + nspans += 1; npix += n + metas.append(dict(name=f"{P}{'u' if x0 else ''}", p=P, x0=x0, off=off, + len=len(blob)-off, nspans=nspans, npix=npix, var=5)) + +# v6: one config per chain depth, so the fit sees spans from 24 to 264 pixels. +for units in range(1, UNITS+1): + P = units * UNITPX + off = len(blob) + nspans = npix = 0 + rows = [] + for y in range(H): + x = 0 + while x < W: + rows.append((y, x)); x += P + blob += struct.pack(">H", len(rows)) + for y, x in rows: + blob += struct.pack(">IH", GVRAM + (YOFF+y)*STRIDE + x*2, + (UNITS-units)*UNITSZ) + # Pad the last span of a row past the visible width; the overrun lands + # in the undisplayed half of the line. + px = np.concatenate([idx[y, x:x+P], np.zeros(max(0, x+P-W), np.uint8)]) + blob += px.astype(">u2").tobytes() + nspans += 1; npix += P + metas.append(dict(name=f"{P}", p=P, x0=0, off=off, len=len(blob)-off, + nspans=nspans, npix=npix, var=6)) + +open(OUT, "wb").write(blob) +with open(META, "w") as f: + f.write("-- generated by tools/bench/prep_spans.py -- do not edit\nreturn {\n") + f.write(f" W={W}, H={H}, total={len(blob)},\n configs = {{\n") + for m in metas: + f.write(" {{var={var}, name=\"{name}\", p={p}, x0={x0}, off={off}," + " len={len}, nspans={nspans}, npix={npix}}},\n".format(**m)) + f.write(" },\n}\n") + +print(f"{SRC} {W}x{H} -> {OUT} {len(blob)} B, {len(metas)} configs") +for m in metas: + print(f" v{m['var']} span {m['name']:>4} px: {m['nspans']:6d} spans, " + f"{m['npix']:6d} px, {m['len']:7d} B " + f"(+{100*m['len']/(2*W*H)-100:.1f}% over bare pixels)") diff --git a/tools/bench/span.lua b/tools/bench/span.lua new file mode 100644 index 0000000..4929d90 --- /dev/null +++ b/tools/bench/span.lua @@ -0,0 +1,218 @@ +-- Measure the cost of a row-linear literal SPAN on the 68000 (FINDINGS 29.5.1). +-- +-- FINDINGS 29 proposes one new decoder mode and prices it at +-- 4 * (50 + 4L*9.08) cycles for a run of L blocks +-- then labels the whole section DERIVED, NOT MEASURED, because both terms are +-- extrapolations: the 50-cycle per-span overhead is hand-derived, and the 9.08 +-- cycles/pixel is a FINDINGS 24 measurement taken at FULL ROW WIDTH with +-- 12-register bursts. A 4-pixel span cannot burst at all. Everything session +-- 8 wants to do downstream optimises over the mode set this number decides, so +-- it goes first. +-- +-- Method: v5 in tools/bench/blit.s walks a stream of per-row span records and +-- copies each span into GVRAM. tools/bench/prep_spans.py emits one stream per +-- span length, every one covering the same whole frame, so the work differs +-- only in how finely it is cut. Regressing +-- cycles = A*spans + B*pixels +-- over the set reads A (the per-span overhead) and B (the per-pixel cost) +-- straight off, and every config also draws a verifiable picture: the frame is +-- cleared before each run and snapshotted after, so a config that timed fast +-- by not writing pixels fails tools/bench/verify_frame256.py. +-- +-- MEASUREMENT SCOPE, unchanged from blit.lua: MAME's gvram_w carries no timing, +-- so these are 68000 instruction cycles against zero-wait-state memory -- a +-- LOWER BOUND on real hardware. Interrupts are masked (SR=$2700). + +M = manager.machine +SP = M.devices[":maincpu"].spaces["program"] + +local function findfile(n) + for _,p in ipairs{"../tools/bench/"..n, "tools/bench/"..n, n} do + local f = io.open(p,"rb"); if f then f:close(); return p end + end + error(n.." not found") +end +local MODE = loadfile(findfile("crtc_mode.lua"))() +local SPEC = loadfile("spans_meta.lua")() + +local FLAG, VAR, ITER, SPTR = 0x18000, 0x18004, 0x18008, 0x1800C +local STREAM = 0x90000 +local GVRAM, GPAL = 0xC00000, 0xE82000 +local CPUHZ = 10000000 -- x68k.cpp:1133, 40_MHz_XTAL/4 +local FRAME12 = CPUHZ / 12 + +local code do local f=io.open("blit.bin","rb"); code=f:read("a"); f:close() end +local blob do local f=io.open("spans.bin","rb"); blob=f:read("a"); f:close() end +local frame do local f=io.open("frame256.bin","rb"); frame=f:read("a"); f:close() end + +local function B(i) return string.byte(frame,i) end +local W, H = B(5)*256+B(6), B(7)*256+B(8) +local PAL0 = 9 +local YOFF = (MODE.height - H) // 2 + +-- Identical packing to blit.lua / show_frame256.lua: shared LSB I per entry. +local function pal6(v) return ((v<<2)|(v>>4)) & 0xff end +local function pack(r,g,b) + local f = {r>>3, g>>3, b>>3} + local best, bestI = nil, 1 + for I = 0,1 do + local e = 0 + for c = 1,3 do + local want = ({r,g,b})[c] + local d = pal6((f[c]<<1)|I) - want + e = e + d*d + end + if best == nil or e < best then best, bestI = e, I end + end + return (f[2]<<11)|(f[1]<<6)|(f[3]<<1)|bestI +end + +local function T() local t=M.time; return t.seconds + t.attoseconds/1e18 end +local function P(s) print("[SPAN] "..s) end + +-- 148 KB one byte at a time is 148k Lua->C calls; longwords cut that by four. +local function push(addr, s, from, len) + local i, n = from, len + while n >= 4 do + SP:write_u32(addr, (string.unpack(">I4", s, i))) + addr, i, n = addr+4, i+4, n-4 + end + while n > 0 do + SP:write_u8(addr, string.byte(s,i)); addr, i, n = addr+1, i+1, n-1 + end +end + +local function clear_picture() -- so a config that writes nothing is caught + for y = YOFF, YOFF+H-1 do + local base = GVRAM + y*1024 + for x = 0, MODE.width-1, 2 do SP:write_u32(base + x*2, 0) end + end +end + +local function setup() + MODE.apply(SP) + for y = 0, MODE.height-1 do + local base = GVRAM + y*1024 + for x = 0, MODE.width-1, 2 do SP:write_u32(base + x*2, 0) end + end + for c = 0, 255 do + local o = PAL0 + c*3 + SP:write_u16(GPAL + c*2, pack(B(o), B(o+1), B(o+2))) + end + for i = 1, #code do SP:write_u8(0x10000+i-1, string.byte(code,i)) end + P(string.format("loaded blit.bin=%d B, %d span configs, picture %dx%d at yoff=%d", + #code, #SPEC.configs, W, H, YOFF)) +end + +local function launch(cfg) + push(STREAM, blob, cfg.off+1, cfg.len) + clear_picture() + -- ~4 emulated seconds per config: 1/55.46 s granularity costs under 0.5%. + local est = cfg.nspans*(cfg.var == 6 and 50 or 60) + cfg.npix*10 + cfg.iter = math.max(4, math.floor(4*CPUHZ/est)) + SP:write_u32(FLAG, 0) + SP:write_u32(VAR, cfg.var) + SP:write_u32(ITER, cfg.iter) + SP:write_u32(SPTR, STREAM) + local cpu = M.devices[":maincpu"] + cpu.state["SR"].value = 0x2700 + cpu.state["SP"].value = 0x8000 + cpu.state["PC"].value = 0x10000 +end + +local results = {} +local function report(cfg, dt) + local cyc = dt * CPUHZ / cfg.iter + results[#results+1] = {cfg=cfg, cyc=cyc} + P(string.format("v%d span %4s px: %5d spans %6d px %d iter in %.4f s -> %8.0f cyc/frame" + .." %5.2f cyc/px %5.1f%% of a 12fps frame", + cfg.var, cfg.name, cfg.nspans, cfg.npix, cfg.iter, dt, cyc, + cyc/cfg.npix, 100*cyc/FRAME12)) +end + +-- Ordinary least squares on cycles = A*spans + B*pixels, no intercept: the +-- 192 row headers and the outer loop are the only work not attributable to a +-- span or a pixel, and at ~10 cycles a row they are 0.2% of the smallest run. +local function fit(rs) + local ss,sp,pp,sy,py = 0,0,0,0,0 + for _,r in ipairs(rs) do + local s,p,y = r.cfg.nspans, r.cfg.npix, r.cyc + ss=ss+s*s; sp=sp+s*p; pp=pp+p*p; sy=sy+s*y; py=py+p*y + end + local det = ss*pp - sp*sp + return (sy*pp - py*sp)/det, (ss*py - sp*sy)/det +end + +local step, st, t0 = 0, "boot", nil + +SUB = emu.add_machine_frame_notifier(function() + local ok, err = pcall(function() + local t = T() + if st == "boot" then + if t < 3.0 then return end + setup(); step = 1; launch(SPEC.configs[1]); st, t0 = "running", nil; return + end + if st == "running" then + local fl = SP:read_u32(FLAG) + if fl == 1 and not t0 then t0 = t; return end + if fl == 0xFF then + report(SPEC.configs[step], t - (t0 or t)) + st = "snap"; return + end + if t > 300 then P("TIMEOUT flag="..string.format("%08X",fl)); M:exit() end + return + end + if st == "snap" then + M.video:snapshot() -- verified by tools/bench/span.sh + step = step + 1 + if SPEC.configs[step] then + launch(SPEC.configs[step]); st, t0 = "running", nil + else + st = "finish" + end + return + end + if st == "finish" then + P("---- measured (instruction cycles only; real GVRAM adds wait states) ----") + for _,v in ipairs{5,6} do + local sub = {} + for _,r in ipairs(results) do if r.cfg.var == v then sub[#sub+1] = r end end + -- v5's fit is over its BURSTING configs only (span length a multiple of + -- the 16-pixel burst). Mixing the remainder-path configs in would hide + -- the two costs behind one bad line; they are reported against the fit + -- instead, which is where the remainder shows up as error. + local fitset = {} + for _,r in ipairs(sub) do + if v == 6 or r.cfg.p % 16 == 0 then fitset[#fitset+1] = r end + end + local A, Bp = fit(fitset) + P(string.format("-- v%d: cycles = %.1f per span + %.3f per pixel" + .." (fitted on %d of %d configs)", v, A, Bp, #fitset, #sub)) + for _,r in ipairs(sub) do + local model = A*r.cfg.nspans + Bp*r.cfg.npix + P(string.format(" span %4s px %8.0f cyc %5.2f cyc/px %6.1f cyc/span" + .." vs fit %+6.1f%%", r.cfg.name, r.cyc, + r.cyc/r.cfg.npix, r.cyc/r.cfg.nspans, 100*(model/r.cyc-1))) + end + -- What the mode decision actually needs: a run of L horizontally + -- adjacent 4x4 blocks is 4 spans of 4L pixels, one per pixel row, and + -- v6 pads each to a whole 24-pixel chain unit. + local line = " -> cycles per 4x4 block in a run of L blocks: " + for _,L in ipairs{1,2,4,8,16,64} do + local px = 4*L + if v == 6 then px = math.ceil(px/24)*24 end + line = line..string.format("L=%d %.0f ", L, 4*(A + px*Bp)/L) + end + P(line.."(V1 is 299.9)") + if v == 5 then + P(" v5's fit only holds where 4L is a whole number of 16-pixel bursts.") + P(" L=1 and L=2 are extrapolations its own measured spans" + .." contradict: 721 and 482.") + end + end + P(" FINDINGS 29 assumed 50.0 per span + 9.080 per pixel, 4 spans per run") + M:exit() + end + end) + if not ok then print("[SPAN] LUA ERROR: "..tostring(err)); M:exit() end +end) diff --git a/tools/bench/span.sh b/tools/bench/span.sh new file mode 100755 index 0000000..9408633 --- /dev/null +++ b/tools/bench/span.sh @@ -0,0 +1,31 @@ +#!/bin/bash +# Measure the cost of a row-linear literal span on the 68000 (FINDINGS 30). +# ~25 s. Run from the repo root. Needs tmp/frame256.bin (check.sh makes it). +# +# NOT part of check.sh, for the same reason blit.s is not: the output is a wall +# timing, so gating on it would make the green light host-sensitive. What IS +# gated here is correctness -- all 23 configs must draw a pixel-exact frame, +# which is what stops a config timing fast by quietly writing nothing. +set -e +cd "$(dirname "$0")/../.." +[ -f tmp/frame256.bin ] || { echo "need tmp/frame256.bin -- run tools/bench/check.sh"; exit 2; } + +python3 tools/bench/prep_spans.py +tools/vasm/vasmm68k_mot -Fbin -o tmp/blit.bin tools/bench/blit.s > /dev/null +mkdir -p tmp/snap_span +rm -f tmp/snap_span/x68000/*.png +( cd tmp && SDL_VIDEODRIVER=dummy timeout -k 5 1800 mame x68000 -bios ipl10 \ + -ramsize 2M -video soft -window -sound none -nothrottle -plugins \ + -autoboot_script ../tools/bench/span.lua \ + -snapshot_directory ./snap_span -snapview native -seconds_to_run 150 \ + > span.log 2>&1 ) +grep -a "^\[SPAN\]" tmp/span.log + +n=0 +for f in tmp/snap_span/x68000/*.png; do + python3 tools/bench/verify_frame256.py "$f" > /dev/null || { + echo "FAIL: $f is not pixel-exact"; python3 tools/bench/verify_frame256.py "$f"; exit 1; } + n=$((n+1)) +done +[ "$n" -eq 23 ] || { echo "FAIL: $n snapshots, expected 23"; exit 1; } +echo "OK $n/23 span configs drew a pixel-exact frame" diff --git a/tools/bench/verify_frame256.py b/tools/bench/verify_frame256.py index 8eda32b..896ad3e 100644 --- a/tools/bench/verify_frame256.py +++ b/tools/bench/verify_frame256.py @@ -1,7 +1,9 @@ #!/usr/bin/env python3 """Regression test for the 256x256 CRTC mode (docs/FINDINGS 23). -Checks tmp/snap256/x68000/0000.png against tmp/frame256.bin: +Checks a native snapshot (default tmp/snap256/x68000/0000.png, override with +argv[1] -- tools/bench/span.lua verifies twelve of them) against +tmp/frame256.bin: 1. native snapshot is 256x512 -- 256 dots, and 512 active scanlines of a 568-line 31.5kHz raster carrying 256 double-scanned graphics rows 2. double-scan pairing is (1,2),(3,4),... -- MAME halves the ABSOLUTE @@ -16,7 +18,8 @@ import struct, sys import numpy as np from PIL import Image -s = np.asarray(Image.open("tmp/snap256/x68000/0000.png").convert("RGB")).astype(int) +snap = sys.argv[1] if len(sys.argv) > 1 else "tmp/snap256/x68000/0000.png" +s = np.asarray(Image.open(snap).convert("RGB")).astype(int) d = open("tmp/frame256.bin", "rb").read() W, H = struct.unpack(">HH", d[4:8]) pal = np.frombuffer(d[8:8+768], np.uint8).reshape(256, 3).astype(int) @@ -53,7 +56,7 @@ if fail: sys.exit(1) mse = ((act - pal[idx]) ** 2).mean() -print(f"OK 256x512 native, double-scan exact, active {W}x{H} pixel-exact, " +print(f"OK {snap}: 256x512 native, double-scan exact, active {W}x{H} pixel-exact, " f"letterbox true black") print(f" palette ceiling vs 24-bit palettised source: " f"{10*np.log10(255**2/mse):.2f} dB ({(I==0).sum()}/256 entries use I=0)")