Put the frame clock on the 68000, and find that the 12 fps frame does not exist
ROADMAP P3 said "needs MFP timer or VBL" and neither can do it. The MFP's timer clock is 16 MHz/4, its prescalers stop at 200 and its data register is 8 bits, so the slowest tick any single timer can make is 78.125 Hz -- 6.5x faster than a frame -- and 4e6/12 is not an integer, so no setting reaches 12 Hz at all. The raster has no whole divide near 12 either: 4 refreshes is 13.86 fps and 5 is 11.09. tools/analysis/23_frame_clock.py walks all 7x256 timer settings rather than asserting it. src/player/clock.i takes the V-DISP falling edge on MFP GPIP4 -- the start of vertical blanking, which is when a player would present -- and adds fps*VTOTAL per edge to a 16-bit accumulator, emitting a tick at 31,500 and keeping the remainder. The long-run rate is fps*VTOTAL/VTOTAL = 12.000000 fps exactly, and both constants are read out of the CRTC at init, so the clock is derived from the registers that generate the raster it counts. Measured over 3,000 refreshes: 3,000 interrupts, 649 ticks where 649.1429 were due. It costs 181.35 clocks per V-DISP, 838 per frame, 0.1006% of the budget -- timed by the 68000 itself, because the host's granularity is 17.64 ms and the interrupt is microseconds. The loop's own cost was calibrated rather than looked up and landed on 38.000002 clocks, which both licenses the subtraction and confirms buscost.py's model; the 181.35 then decomposes exactly, leaving 43.99 clocks for the interrupt exception -- the textbook 44, measured. THE ONE THAT MOVES SOMETHING: 12 fps on a 55.4577 Hz raster is 4.6215 refreshes, so a frame is shown for 4 refreshes (72.13 ms) or 5 (90.16 ms), 37.9% of them short. The 833,333-clock budget every figure in this project is priced against is the MEAN slot, and the short one is 13.4% under it. The cadence was already in the tree unnamed: stream.lua's tick is sampled at frame boundaries, so its gaps were always 4 or 5, and every host-paced result in FINDINGS 49/51 carried it. P3 moved who produces it onto the machine and made it visible. It is not a dropped frame -- the pace gate lets an overrun eat the next frame's idle -- and on the gate container it costs 4 frames of 120 their idle against 1 for the nominal model, most of that the frame-0 transient at 111% of budget. stream.s counts it now, and the rig matches an offline model of the divider exactly. Also struck: MAME's raster runs 2.22% fast. refresh_mode() builds the frame period from scr.max_x*scr.max_y with scr.max_x = m_htotal - 8, one character cell short and an inclusive bound used as a count, so it runs at 56.6901 Hz where the registers say 55.4577 -- agreeing to six digits with the arithmetic. Every "1/55.46 s granularity" note in this tree was wrong and is 1/56.69 s, corrected in six files with the derivation put once in crtc_mode.lua. No conclusion changes and no 68000 cycle figure moves; the CPU clock is unrelated to the screen. But anything paced by the raster runs fast under MAME, so the rig reports both rates and prices the interrupt against the hardware's. decode.s and frame.i are unchanged; decode.bin is still 1,296 B at the same MD5. The pace gate's wait loop is byte-for-byte the one FINDINGS 51 measured and the free-running path executes none of the new code. check.sh gains two stages: the clock's own measurement, and 120 frames decoded pixel-exact with nothing outside the machine deciding when a frame may start. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -19,7 +19,9 @@ of its own snapshot, 2x nearest-neighbour, no filtering.
|
||||
**The player, running.** 119 frames out of a **256 KB ring buffer on an emulated
|
||||
stock 2 MB X68000**, paced to a 12 fps frame clock, streamed from a host file at
|
||||
488 KB/s by `src/player/stream.s` with no Lua in the decode path. Source on the
|
||||
left, the machine's screen on the right.
|
||||
left, the machine's screen on the right. (This recording was paced by the host;
|
||||
the 68000 now keeps that clock itself, off the CRTC's V-DISP, and the same 120
|
||||
frames decode pixel-exact under it — `src/player/clock.i`, FINDINGS 54.)
|
||||
|
||||
<video src="docs/img/player.webm" controls muted loop width="100%"></video>
|
||||
|
||||
|
||||
+198
-2
@@ -4535,8 +4535,10 @@ because it is a pure copy, and it still runs alone.
|
||||
**The stages are exactly additive on the exact core.** `P1 + P2 - SCENE = 330`
|
||||
clocks, and `tables + P1 + P2 - 2x330 = 253,614 = BOOT`, to the clock — 330 is
|
||||
the front-end's own per-pass overhead. On MAME the same identity closes to 1.8%,
|
||||
which is one tick of its 1/55.46 s clock over the 0.99 s run. Two instruments,
|
||||
two granularities, one arithmetic.
|
||||
which is one tick of its host clock over the 0.99 s run. Two instruments,
|
||||
two granularities, one arithmetic. (That tick was written here as 1/55.46 s and
|
||||
is **1/56.69 s** — MAME's raster, not the hardware's, 54.5. 17.64 ms over 990 ms
|
||||
is 1.78%, so the sentence was right and the label was wrong.)
|
||||
|
||||
### 53.3 The scratch tables are scene-independent, so they are not in the scene path
|
||||
|
||||
@@ -4618,3 +4620,197 @@ black with `I = 0`. That half of P2 is encoder-side, it changes the container,
|
||||
and it moves every constant fitted to the gate container, so it is a re-encode
|
||||
plus a re-measurement rather than an edit. `load.i` is ready for it: it reads
|
||||
whatever the palette section holds and reports the darkest index either way.
|
||||
|
||||
---
|
||||
|
||||
## 54. The frame clock moves onto the 68000, and the 12 fps frame turns out never to have existed (session 22)
|
||||
|
||||
ROADMAP P3, and the item was phrased "needs MFP timer or VBL" — which quietly
|
||||
assumes one of those can do it. Neither can, and finding out why produced a
|
||||
better clock than either and a correction to an instrument the whole tree reads.
|
||||
|
||||
`src/player/clock.i` is the clock; `src/player/clockgate.s` and
|
||||
`tools/bench/clock.lua` measure it; `tools/analysis/23_frame_clock.py`
|
||||
enumerates the space it was chosen from and prices its cadence. Two new stages
|
||||
in `tools/bench/check.sh` gate it.
|
||||
|
||||
**Layer: MAME 0.277's emulated X68000, not real hardware.** The MFP, the CRTC
|
||||
and the interrupt sequence are all the emulator's. Where the emulator and the
|
||||
registers disagree — and they do, 54.5 — the code is built on the registers.
|
||||
|
||||
### 54.1 No MFP timer can tick at 12 Hz, and none can tick as slowly as a frame
|
||||
|
||||
The MC68901's timer clock on this board is 16 MHz / 4 = 4 MHz
|
||||
(`sharp/x68k.cpp:1027-1028`), its prescaler ladder is `{4, 10, 16, 50, 64, 100,
|
||||
200}` (`machine/mc68901.cpp:173`) and its data register is 8 bits. So:
|
||||
|
||||
* the **slowest** tick a single timer can produce is 4e6/(200·256) =
|
||||
**78.125 Hz**, which is 6.5× faster than a 12 fps frame — a software divider
|
||||
is required whatever the source;
|
||||
* **4e6/12 = 333,333.33 is not an integer**, so no prescale/data pair divides to
|
||||
12 Hz at all. `23_frame_clock.py` walks all 7 × 256 of them and finds zero.
|
||||
|
||||
A timer clock is therefore not "the simple option". It is a divider *plus* an
|
||||
interrupt rate 3.6× higher than the raster's, at an arbitrary phase against the
|
||||
scan.
|
||||
|
||||
### 54.2 The raster cannot do it by whole division either — and the fix is exact
|
||||
|
||||
V-DISP is on MFP GPIP4 (`x68k.cpp:1139`), and the same pin is Timer A's event
|
||||
input (`mc68901.cpp:167`, `GPIO_TIMER = {GPIP_4, GPIP_3}`); its interrupt is
|
||||
channel 6, `IR_GPIP_4 = $40` (`mc68901.cpp:76`). The raster is 31,500/568 =
|
||||
**55.4577 Hz** exactly. Every whole divide misses:
|
||||
|
||||
| Timer A event count | fps | error |
|
||||
|---:|---:|---:|
|
||||
| 4 | 13.8644 | +15.54% |
|
||||
| 5 | 11.0915 | −7.57% |
|
||||
|
||||
12 fps needs **4.6215 refreshes per frame**. So the divider keeps a remainder:
|
||||
|
||||
```
|
||||
each V-DISP: acc += fps*VTOTAL ; 12*568 = 6816
|
||||
if acc >= 31500: acc -= 31500 ; PACE += 1
|
||||
```
|
||||
|
||||
Long-run rate is `fps·VTOTAL/VTOTAL` = **12.000000 fps exactly**, with a
|
||||
remainder that never accumulates. Both constants are **read out of the CRTC at
|
||||
init** — `R04+1` for VTOTAL, `R20` bit 4 checked for the 31.5 kHz mode — so the
|
||||
clock is derived from the registers that generate the raster it counts, and the
|
||||
two cannot drift apart. The accumulator peaks at 38,316, so it is 16-bit
|
||||
arithmetic on a 68000; `clk_init` refuses rather than overflow (the ceiling is
|
||||
fps < 59.9 at this VTOTAL).
|
||||
|
||||
**Measured over 3,000 refreshes: 3,000 interrupts, 649 ticks, where 649.1429
|
||||
were exactly due — an error of −0.14 ticks, i.e. the remainder still held.** The
|
||||
gate is stated in ticks and not in ppm on purpose: a remainder-keeping divider
|
||||
is off by at most one tick over *any* window, so quoting ppm would let a longer
|
||||
window advertise a tighter clock for nothing.
|
||||
|
||||
### 54.3 It costs 181.35 clocks per V-DISP — 838 per frame, 0.10% of the budget
|
||||
|
||||
The host cannot time this: MAME's Lua sees the machine once per screen frame,
|
||||
17.64 ms, and the interrupt costs microseconds. So the **68000 times it
|
||||
itself**. `clockgate.s` runs a one-instruction loop for a window of 3,000
|
||||
refreshes with the clock off and again with it armed:
|
||||
|
||||
```
|
||||
clock off: iters0·L = clocks in the window -> L
|
||||
clock on: iters1·L + ints·H = clocks in the window -> H
|
||||
```
|
||||
|
||||
`L` is **calibrated, not looked up** — the point is to price the clock on the
|
||||
machine rather than against `buscost.py`, which is the table being checked.
|
||||
|
||||
| | |
|
||||
|---|---:|
|
||||
| loop iteration, calibrated over 13,926,121 of them | **38.000002 clocks** |
|
||||
| per V-DISP interrupt, over 3,000 | **181.35 clocks** |
|
||||
| per 12 fps frame (4.6215 interrupts) | **838 clocks = 0.1006%** |
|
||||
|
||||
`L` landing on a whole number to seven digits is the check that licenses the
|
||||
subtraction, and it is also an independent confirmation of `buscost.py`'s model:
|
||||
`addq.l #1,(xxx).L` is 3 instruction words + 4 data accesses = 7 bus cycles = 28
|
||||
clocks, plus 10 for the `bra.s`.
|
||||
|
||||
**The 181.35 decomposes exactly.** By the same model the handler body is 130
|
||||
clocks on a V-DISP that emits no tick and 164 on one that does; over the measured
|
||||
649/3,000 mix that is 137.355, leaving **43.99 clocks for the interrupt exception
|
||||
sequence** — the textbook 44, measured here rather than recalled.
|
||||
|
||||
For comparison, the cheapest exact MFP-timer clock would interrupt 16.7 times a
|
||||
frame instead of 4.6: **3.6× the cost, for a tick with no fixed relationship to
|
||||
the scan.**
|
||||
|
||||
### 54.4 THE ONE THAT MOVES SOMETHING: there is no 83.33 ms frame, and there never was
|
||||
|
||||
12 fps on a 55.4577 Hz raster is 4.6215 refreshes, so a frame is shown for **4
|
||||
refreshes (72.13 ms) or 5 (90.16 ms)** — 37.9% of them short. The 833,333-clock
|
||||
budget every figure in this project is priced against is the **mean** slot, and
|
||||
the short slot is **13.4% under it**.
|
||||
|
||||
With the per-frame decode costs (`tmp/c68k_frames.csv`, 120 frames of the gate
|
||||
container) run through the actual divider and the actual pace gate:
|
||||
|
||||
| tick source | short slot | frames over it | no idle left |
|
||||
|---|---:|---:|---:|
|
||||
| nominal 1/fps model (no raster has it) | 833,333 | 1/120 | **1/120** |
|
||||
| the host tick, as `stream.lua` really emits it | 705,590 | 10/120 | **4/120** |
|
||||
| the 68000's clock, hardware raster | 721,270 | 10/120 | **4/120** |
|
||||
| the 68000's clock, MAME's raster | 705,590 | 10/120 | **4/120** |
|
||||
|
||||
**The cadence was already there and nothing had named it.** `stream.lua`'s tick
|
||||
is `floor((t - t_rel) * fps)` — which *looks* uniform and is not, because Lua
|
||||
only sees the machine at frame boundaries, so its ticks land on refreshes and its
|
||||
gaps are the same two whole numbers. Every host-paced result in FINDINGS 49 and
|
||||
51 already carried a 4/5 cadence. **P3 did not introduce it. It moved who
|
||||
produces it onto the machine and made it visible.**
|
||||
|
||||
**A short slot is not a dropped frame.** The pace gate says only "not before
|
||||
tick i", so a frame that overruns spends the next frame's idle and the clock
|
||||
recovers itself; the cost is one frame presented a refresh late. What the table
|
||||
counts is frames with no idle left, and the difference between the nominal row
|
||||
and the raster rows — **1 against 4** — is the entire price of the cadence on
|
||||
this container.
|
||||
|
||||
**The expensive frame is frame 0, at 923,146 clocks = 111% of the nominal
|
||||
budget**: the first frame of a scene has nothing to SKIP against, so it is the
|
||||
whole picture in one slot. Most of what follows it in those counts is that
|
||||
transient draining. It also means the cost lands **at a scene change**, next to
|
||||
FINDINGS 53.2's 18.96 ms of loader and the seek — not spread over the window.
|
||||
|
||||
`src/player/stream.s` now counts this itself (`LATEFR`/`LATEMAX`/`LATE1ST`), and
|
||||
the rig's count matches the offline model **exactly**: 4/120, first at frame 1,
|
||||
on both tick sources. The counter sits ahead of the wait loop and the
|
||||
free-running path executes none of it, so FINDINGS 49's figures are untouched.
|
||||
|
||||
### 54.5 MAME's raster runs 2.22% fast, and the whole tree has been sampling it
|
||||
|
||||
`x68k_crtc.cpp refresh_mode()` computes the frame period as
|
||||
`(scr.max_x * scr.max_y)` dots over the dot clock, with
|
||||
`scr.max_x = m_htotal - 8` — one character cell short, and an **inclusive
|
||||
rectangle bound used as a count**. In the 256-wide mode that is 360 where the
|
||||
registers say 368, so MAME's refresh is fast by **368/360 = 1.02222**:
|
||||
|
||||
* registers: 31,500/568 = **55.4577 Hz**
|
||||
* MAME, measured by `clock.lua` over 3,000 frames: **56.6901 Hz**
|
||||
|
||||
The two agree to six digits with `clock_69m()/6 / (360·568)`, so this is the
|
||||
mechanism and not a coincidence. Consequences, and the third one is why it is
|
||||
worth this much space:
|
||||
|
||||
1. **Every "1/55.46 s granularity" note in this tree was wrong** — it is
|
||||
1/56.69 s, 17.64 ms. Corrected in `decode.lua`, `load.lua`, `span.lua`,
|
||||
`blit.lua`, `loadgate.s` and `check.sh`, with the derivation put once in
|
||||
`crtc_mode.lua`. No conclusion changes: 53.2's "one tick over the 0.99 s run"
|
||||
is 1.78% at the corrected figure and was quoted as 1.8%.
|
||||
2. **68000 cycle figures are untouched.** The CPU clock is 40 MHz/4 and has
|
||||
nothing to do with the screen. Nothing in FINDINGS 24–53 moves.
|
||||
3. **A raster-paced player runs 2.22% fast under MAME**, so the rig measures
|
||||
12.267 fps where the hardware would give 12.000. `clock.lua` reports both and
|
||||
de-skews, and `clock_run.sh` prices the interrupt against the **hardware**
|
||||
refresh count — charging the player the emulator's extra interrupts would
|
||||
overstate the cost by that same 2.2%.
|
||||
|
||||
**Do not "fix" 55.4577 to match the measurement.** It is the hardware's, derived
|
||||
from the dot clocks, and it is what the divider is built on.
|
||||
|
||||
### 54.6 What had to be turned off, and why it is in the file
|
||||
|
||||
The rigs launch the 68000 at `SR=$2700` into a machine the IPL ROM has already
|
||||
booted, so the MFP arrives with whatever IOCS enabled on it and vectors pointing
|
||||
into IOCS. Lowering the mask without disarming it would vector into code we did
|
||||
not put there. `clk_init` writes `IERA = IERB = 0` first — which on the MC68901
|
||||
clears the matching pending bits with them (`mc68901.cpp REGISTER_IERA/B`,
|
||||
`m_ipr &= m_ier`) — then arms GPIP4 alone, takes the falling edge (AER bit 4
|
||||
clear: the **start of vertical blanking**, which is when a player would present),
|
||||
and drops to `SR=$2500`. Levels 1–5 stay masked, so the DMAC (IRQ3) and the SCC
|
||||
(IRQ5) cannot get in. `VR` is written with **S clear**, so an acknowledge clears
|
||||
the pending bit by itself and the handler needs no end-of-interrupt write.
|
||||
|
||||
The handler saves only the **low word** of `d0`, because every operation in it is
|
||||
a word operation — which is legal precisely because `clk_init` proved the
|
||||
accumulator fits 16 bits. `decode.s` and `frame.i` were checked for stack tricks
|
||||
before the mask was lowered: the only `a7` use in either is one `move.l a1,-(sp)`
|
||||
pair, so an interrupt cannot corrupt decoder state. The 120-frame self-paced
|
||||
decode being pixel-exact is the test of that, and it is gated.
|
||||
|
||||
+32
-4
@@ -2,6 +2,7 @@
|
||||
|
||||
Written end of session 19 (2026-08-24), against a tree that is ALL GREEN.
|
||||
Amended end of session 21: P1 done, P2 half done (FINDINGS 53).
|
||||
Amended end of session 22: P3 done (FINDINGS 54).
|
||||
|
||||
**THE COMPLETION TARGET IS M3, THE VERTICAL SLICE** (USER DECISION): one scene
|
||||
tree — a decision point, two outcomes, a death clip — with audio, streaming from
|
||||
@@ -124,10 +125,37 @@ re-measurement rather than an edit. Until then the letterbox gets the palette's
|
||||
closest thing to black (index 255 on the gate container); `load.i` reports
|
||||
whichever index that is and needs no change when it becomes 0.
|
||||
|
||||
**P3. A real frame clock.** `stream.s` has `PACE`/`PACEON` (`$18034`/`$18038`)
|
||||
but the 12 fps tick comes from the Lua producer. Needs MFP timer or VBL. Keep
|
||||
`PACEON=0` free-run working — the wrap gate uses it and every FINDINGS 49 figure
|
||||
depends on it.
|
||||
~~**P3. A real frame clock.**~~ **DONE, session 22 — FINDINGS 54.**
|
||||
`src/player/clock.i` derives the tick from the CRTC's own V-DISP through the
|
||||
MFP, with a remainder-keeping divider whose two constants are read out of the
|
||||
CRTC at init. **Exactly 12.000000 fps, by construction** — measured at 649 ticks
|
||||
over 3,000 refreshes where 649.1429 were due, so the remainder still held and
|
||||
nothing accumulated. **181.35 clocks per V-DISP, 838 per frame, 0.1006% of the
|
||||
budget**, timed by the 68000 itself because the host's 17.64 ms granularity
|
||||
cannot see it. `PACEON=0` free-run is untouched and so is the wait loop; the
|
||||
free-running path executes none of the new code.
|
||||
|
||||
The item said "MFP timer or VBL" and **neither can do it alone**: 4e6/12 is not
|
||||
an integer and no prescale/data pair reaches 12 Hz, while the slowest MFP tick
|
||||
of any kind is 78.125 Hz; and the raster's 55.4577 Hz has no whole divide near
|
||||
12 either (4 gives 13.86, 5 gives 11.09). `tools/analysis/23_frame_clock.py`
|
||||
walks the whole space rather than asserting it.
|
||||
|
||||
**What it exposed is bigger than the item.** 12 fps on a 55.4577 Hz raster is
|
||||
4.6215 refreshes, so a frame gets **4 refreshes (72.13 ms) or 5 (90.16 ms)** and
|
||||
**there is no 83.33 ms frame** — that figure is the mean slot, and 37.9% of slots
|
||||
are 13.4% under it. The cadence was ALREADY in every host-paced result in
|
||||
FINDINGS 49/51, because `stream.lua`'s tick is sampled at frame boundaries and
|
||||
its gaps were always 4 or 5; nothing had named it. On the gate container it
|
||||
costs 4 frames of 120 their idle against 1 for the nominal model. **It is not a
|
||||
dropped frame** — the pace gate lets an overrun eat the next frame's idle and
|
||||
the clock recovers — but it means every budget in this project is priced against
|
||||
a slot 37.9% of frames do not get. 54.4.
|
||||
|
||||
**Also struck: MAME's raster runs 2.22% fast** (`refresh_mode()` builds the frame
|
||||
period from `htotal - 8`), so the tree's "1/55.46 s granularity" was 1/56.69 s
|
||||
throughout. No 68000 cycle figure moves — the CPU clock is unrelated to the
|
||||
screen — but anything paced by the raster does. 54.5.
|
||||
|
||||
**P4. Real transport.** Drive the MB89352 instead of a host file. The `W`
|
||||
handshake — clocks stolen per delivered byte, bracketed 5..12 by MC68450 Fig
|
||||
|
||||
+116
-1
@@ -1,3 +1,117 @@
|
||||
# Status & next-session handoff — end of session 22 (2026-08-24)
|
||||
|
||||
## Session 22: the frame clock moves onto the 68000, and the 12 fps frame turns out not to exist
|
||||
|
||||
**Green light first and last: `./tools/bench/check.sh` was ALL GREEN before any
|
||||
of this and ALL GREEN after**, 120/120 on both cores, no `TRUNCATED`, plus two
|
||||
new frame-clock stages.
|
||||
|
||||
**ROADMAP P3 is DONE. FINDINGS 54.** P3 was one of the two items session 21's
|
||||
handoff called buildable here, and it is the one that touches every other number
|
||||
in the project — because the tick is what the word "frame" in "% of a frame"
|
||||
means.
|
||||
|
||||
**1. Neither of the two sources P3 named can do it, and the enumeration is the
|
||||
finding.** The MFP's timer clock is 16 MHz/4 = 4 MHz, its prescalers are
|
||||
`{4,10,16,50,64,100,200}` and its data register is 8 bits, so the **slowest tick
|
||||
any single timer can make is 78.125 Hz** — 6.5x faster than a frame — and
|
||||
**4e6/12 = 333,333.33 is not an integer**, so no setting reaches 12 Hz at all.
|
||||
The raster has no whole divide near 12 either: 4 refreshes is 13.86 fps and 5 is
|
||||
11.09. `tools/analysis/23_frame_clock.py` walks all 7x256 timer settings rather
|
||||
than asserting it. 54.1, 54.2.
|
||||
|
||||
**2. The clock is the raster with a remainder, and it is exact by construction.**
|
||||
`src/player/clock.i` takes the V-DISP falling edge on MFP GPIP4 — the start of
|
||||
vertical blanking, which is when a player would present — and adds `fps*VTOTAL`
|
||||
per edge to a 16-bit accumulator, emitting a tick at 31,500 and keeping the
|
||||
remainder. Long-run rate is `fps*VTOTAL/VTOTAL` = **12.000000 fps exactly**.
|
||||
Both constants are **read out of the CRTC at init**, so the clock is derived from
|
||||
the registers that generate the raster it counts. Measured: **3,000 interrupts,
|
||||
649 ticks where 649.1429 were due**. The gate is stated in ticks, not ppm — a
|
||||
remainder is off by at most one tick over any window, so ppm would let a longer
|
||||
window advertise a tighter clock for free. 54.2.
|
||||
|
||||
**3. It costs 181.35 clocks per V-DISP; 838 per frame; 0.1006% of the budget.**
|
||||
The host cannot time this — its granularity is 17.64 ms and the interrupt is
|
||||
microseconds — so **the 68000 times it itself**: a one-instruction loop over a
|
||||
3,000-refresh window, run with the clock off and on, with the loop's own cost
|
||||
calibrated rather than looked up. The calibration landed on **38.000002 clocks**
|
||||
per iteration, which is both the check that licenses the subtraction and an
|
||||
independent confirmation of `buscost.py`'s model. The 181.35 then decomposes
|
||||
exactly, leaving **43.99 clocks for the interrupt exception** — the textbook 44,
|
||||
measured rather than recalled. A timer-based clock would have cost 3.6x this at
|
||||
an arbitrary phase against the scan. 54.3.
|
||||
|
||||
**4. THE ONE THAT MOVES SOMETHING: there is no 83.33 ms frame, and there never
|
||||
was.** 12 fps on a 55.4577 Hz raster is 4.6215 refreshes, so a frame is shown for
|
||||
**4 refreshes (72.13 ms) or 5 (90.16 ms)**, 37.9% of them short. The
|
||||
833,333-clock budget every figure in this project is priced against is the
|
||||
**mean** slot; the short one is **13.4% under it**, and 10 of the gate
|
||||
container's 120 frames do not fit it.
|
||||
|
||||
**And the cadence was already in the tree, unnamed.** `stream.lua`'s tick is
|
||||
`floor((t - t_rel) * fps)`, which looks uniform and is not: Lua only sees the
|
||||
machine at frame boundaries, so its ticks land on refreshes and its gaps were
|
||||
always 4 or 5. **Every host-paced result in FINDINGS 49 and 51 already carried
|
||||
this cadence.** P3 did not introduce it; it moved who produces it onto the
|
||||
machine and made it visible.
|
||||
|
||||
**It is not a dropped frame.** The pace gate says only "not before tick i", so an
|
||||
overrun eats the next frame's idle and the clock recovers itself; the cost is one
|
||||
frame presented a refresh late. On the gate container that is **4 frames of 120
|
||||
with no idle left, against 1 for the nominal model** — and the expensive one is
|
||||
**frame 0 at 111% of budget**, because the first frame of a scene has nothing to
|
||||
SKIP against. So the cost lands **at a scene change**, next to 53.2's 18.96 ms of
|
||||
loader and the seek. `stream.s` counts this itself now, and the rig's count
|
||||
matches an offline model of the divider **exactly**: 4/120, first at frame 1, on
|
||||
both tick sources. 54.4.
|
||||
|
||||
**5. An instrument correction the whole tree was reading.** `x68k_crtc.cpp
|
||||
refresh_mode()` builds the frame period from `scr.max_x * scr.max_y` with
|
||||
`scr.max_x = m_htotal - 8` — one character cell short, an inclusive bound used as
|
||||
a count. **MAME's raster is fast by 368/360 = 2.2222%**: 56.6901 Hz measured
|
||||
against the registers' 55.4577, agreeing to six digits with the arithmetic. So
|
||||
every "1/55.46 s granularity" note in this tree was wrong and is **1/56.69 s**;
|
||||
corrected in six files with the derivation put once in `crtc_mode.lua`. **No
|
||||
conclusion changes and no 68000 cycle figure moves** — the CPU clock is unrelated
|
||||
to the screen — but anything *paced* by the raster runs 2.22% fast under MAME, so
|
||||
`clock.lua` reports both rates and de-skews, and the interrupt is priced against
|
||||
the hardware refresh count. 54.5.
|
||||
|
||||
**New in the tree:** `src/player/clock.i` (the clock) and
|
||||
`src/player/clockgate.s` (its measurement front-end); `tools/bench/clock.lua`,
|
||||
`clock_cost.py`, `clock_run.sh` (the rig); `tools/analysis/23_frame_clock.py`
|
||||
(the enumeration and the cadence pricing). `stream.s` gains `CLKON` and a
|
||||
late-frame counter ahead of the wait loop; `stream.lua` gains `DLX_PACE=2` and
|
||||
takes its deadlines from the ticks the machine actually emitted rather than from
|
||||
a host model; `pace_run.sh` gains `DLX_PACE` selection, with the default tag left
|
||||
alone so `pace_sweep.sh` still finds its logs. `check.sh` gains two stages: the
|
||||
clock's own measurement, and 120 frames decoded pixel-exact with nothing outside
|
||||
the machine deciding when a frame may start.
|
||||
|
||||
**`decode.s` and `frame.i` are unchanged.** `decode.bin` is still 1,296 B at the
|
||||
same MD5. The pace gate's wait loop is byte-for-byte the one FINDINGS 51
|
||||
measured, and the free-running path executes none of the new code, so every
|
||||
FINDINGS 49 figure stands.
|
||||
|
||||
**Still open in P2:** unchanged — the encoder does not reserve a black entry
|
||||
(23.4), so the letterbox still gets the palette's closest thing to black.
|
||||
|
||||
**Next:** P5 (per-record index, prefill policy, the accumulated-slack rule in the
|
||||
player rather than the rig) is buildable here and is now the last M2 item that
|
||||
is. G1 (import the scene graph) still needs fetching, and is still the one that
|
||||
would let this tree ask what the worst gap between consecutive decision points
|
||||
is. P4 still decides the project and still cannot be measured here.
|
||||
|
||||
**A question 54.4 raises and does not answer:** every rate-control and budget
|
||||
figure in this project is fitted to an 833,333-clock frame, and 37.9% of frames
|
||||
get 721,270. Whether the encoder should be fitted to the SHORT slot instead of
|
||||
the mean is a re-encode plus a re-measurement — the same class of change as the
|
||||
reserved black entry — and it should be decided with P5's numbers in hand, not
|
||||
before.
|
||||
|
||||
---
|
||||
|
||||
# Status & next-session handoff — end of session 21 (2026-08-24)
|
||||
|
||||
## Session 21: the loader moves onto the 68000, and a scene change gets a price
|
||||
@@ -68,7 +182,8 @@ what this project is short of. **Derived, not measured.** 53.6.
|
||||
gains a `--loadraw` mode, which also makes its flag-watch address a variable
|
||||
rather than a constant. `check.sh` gains a stage that gates byte-exactness on
|
||||
both cores, and deliberately does **not** gate the cycle counts — MAME's clock
|
||||
is 1/55.46 s and a wall timing would make the green light host-sensitive, the
|
||||
is 1/56.69 s (1/55.46 when that was written; 54.5) and a wall timing would make
|
||||
the green light host-sensitive, the
|
||||
same reason `blit.s` and `span.sh` are not in it.
|
||||
|
||||
**`decode.s` and `stream.s` are unchanged.** Nothing in the per-frame path was
|
||||
|
||||
@@ -0,0 +1,202 @@
|
||||
; ---------------------------------------------------------------------------
|
||||
; clock.i -- the FRAME CLOCK, on the 68000 itself. ROADMAP P3.
|
||||
;
|
||||
; src/player/stream.s has a pace gate: frame i may not START before tick i, and
|
||||
; PACE is the tick counter. Until now PACE was written by tools/bench/
|
||||
; stream.lua, i.e. by the host, off the host's idea of what 12 fps means. That
|
||||
; was honest for what FINDINGS 49/51 were measuring -- arrival times against a
|
||||
; deadline -- and it is not a player. A player has no host. These are the
|
||||
; bytes that replace it.
|
||||
;
|
||||
; WHAT THE MACHINE ACTUALLY OFFERS, because "use the MFP timer or vblank" hides
|
||||
; a real constraint. The MC68901's timer clock on this board is 16 MHz / 4 =
|
||||
; 4 MHz (MAME 0.277 src/mame/sharp/x68k.cpp:1027-1028), its prescaler ladder is
|
||||
; {4, 10, 16, 50, 64, 100, 200} (src/devices/machine/mc68901.cpp:173) and its
|
||||
; data register is 8 bits. The SLOWEST tick a single MFP timer can produce is
|
||||
; therefore 4e6 / (200*256) = 78.125 Hz, and 4e6/12 = 333,333.33 is not even an
|
||||
; integer -- so no prescaler/data pair ticks at 12 Hz, and no timer at any
|
||||
; setting ticks as slowly as a 12 fps frame. A frame clock needs a divider in
|
||||
; software whichever source it is built on. tools/analysis/23_frame_clock.py
|
||||
; enumerates the whole space rather than asserting this.
|
||||
;
|
||||
; So the source is the RASTER, and that is a better answer than a timer anyway.
|
||||
; GPIP4 on the MFP is V-DISP (x68k.cpp:1139, `m_crtc->vdisp_cb().set(i4_w)`),
|
||||
; high while the display is active; the same signal is the MFP's Timer A event
|
||||
; input (mc68901.cpp:167, GPIO_TIMER = {GPIP_4, GPIP_3}). Its interrupt is
|
||||
; channel 6, IR_GPIP_4 = $40 in IERB/IPRB/IMRB (mc68901.cpp:76). We take the
|
||||
; FALLING edge (AER bit 4 = 0), which is the start of vertical blanking -- the
|
||||
; instant a player would present a finished frame, so the clock and the flip
|
||||
; are the same event rather than two events with a phase between them.
|
||||
;
|
||||
; THE DIVIDER IS EXACT, AND IT IS EXACT BY CONSTRUCTION. The raster is
|
||||
;
|
||||
; 31,500 lines/s / (R04 + 1) lines/frame
|
||||
;
|
||||
; and 31,500 / 568 = 55.4577 Hz is not a multiple of 12, so a whole-number
|
||||
; divide cannot do it: 4 refreshes is 13.87 fps and 5 is 11.09 fps. Instead
|
||||
; each V-DISP adds `fps * (R04+1)` to an accumulator and a frame tick is emitted
|
||||
; whenever it reaches 31,500, keeping the remainder:
|
||||
;
|
||||
; acc += fps*VTOTAL ; if acc >= HFREQ: acc -= HFREQ ; PACE += 1
|
||||
;
|
||||
; Over VTOTAL/gcd raster frames that emits exactly fps*VTOTAL/gcd ticks, so the
|
||||
; long-run rate is fps*VTOTAL/VTOTAL = fps EXACTLY, with a bounded remainder and
|
||||
; ZERO accumulated drift -- not 12.0001, not 11.9998. It holds for any fps and
|
||||
; any vertical geometry, which is why the two constants are READ OUT OF THE
|
||||
; CRTC at init rather than assembled in: the clock is derived from the same
|
||||
; registers that generate the raster it is counting, so the two cannot disagree.
|
||||
;
|
||||
; WHAT IT COSTS IN CADENCE, WHICH IS THE PART THAT IS NOT FREE. 12 fps on a
|
||||
; 55.4577 Hz raster is 4.6215 refreshes per frame, so a frame is shown for
|
||||
; either 4 or 5 refreshes -- 72.13 ms or 90.16 ms. Nothing can change that;
|
||||
; it is the display's quantisation, not the clock's error, and a timer-derived
|
||||
; clock would have exactly the same cadence with an arbitrary phase against the
|
||||
; raster on top. It does mean the slot a frame gets is NOT always the 83.33 ms
|
||||
; every budget in this project is priced against, and the short slot is 13.4%
|
||||
; under it. tools/analysis/23_frame_clock.py prices that; do not read this file
|
||||
; as a claim that the clock made the budget bigger.
|
||||
;
|
||||
; INTERRUPTS, AND WHAT HAD TO BE TURNED OFF. The rigs launch the 68000 at
|
||||
; SR=$2700 with everything masked, into a machine the IPL ROM has already booted
|
||||
; -- so the MFP arrives with whatever IOCS enabled on it (keyboard receive,
|
||||
; Timer C, its own V-DISP handler) and vectors pointing into IOCS. Lowering the
|
||||
; mask without disarming the MFP would vector into code we did not put there.
|
||||
; clk_init therefore writes IERA = IERB = 0 first, which on the MC68901 also
|
||||
; clears the matching pending bits (mc68901.cpp REGISTER_IERA/B: `m_ipr &=
|
||||
; m_ier`), and only then arms GPIP4 alone. Levels 1-5 stay masked at SR=$2500,
|
||||
; so the DMAC (IRQ3) and the SCC (IRQ5) cannot get in either; level 7 is the
|
||||
; front-panel NMI and is not ours to mask.
|
||||
;
|
||||
; The vector is the MFP's own: VR is written with the S bit CLEAR, so the
|
||||
; in-service register is not used and an acknowledge clears the pending bit by
|
||||
; itself (mc68901.cpp get_vector). No end-of-interrupt write in the handler.
|
||||
; ---------------------------------------------------------------------------
|
||||
|
||||
; --- MFP registers. The device sits on D0-D7 of a 16-bit bus (x68k.cpp:793,
|
||||
; `.umask16(0x00ff)`), so register n is one BYTE at $E88001 + 2n.
|
||||
MFP = $E88001
|
||||
MFP_GPIP = MFP+0*2
|
||||
MFP_AER = MFP+1*2
|
||||
MFP_DDR = MFP+2*2
|
||||
MFP_IERA = MFP+3*2
|
||||
MFP_IERB = MFP+4*2
|
||||
MFP_IPRA = MFP+5*2
|
||||
MFP_IPRB = MFP+6*2
|
||||
MFP_ISRA = MFP+7*2
|
||||
MFP_ISRB = MFP+8*2
|
||||
MFP_IMRA = MFP+9*2
|
||||
MFP_IMRB = MFP+10*2
|
||||
MFP_VR = MFP+11*2
|
||||
|
||||
MFP_GPIP4 = 4 ; bit number of V-DISP in GPIP/AER/DDR
|
||||
MFP_IVDISP = $40 ; IR_GPIP_4, channel 6, in IERB/IPRB/IMRB
|
||||
MFP_VBASE = $40 ; vector base; S clear -> no in-service register
|
||||
CLK_VEC = (MFP_VBASE+6)*4 ; $118: MFP channel 6 vector, as an ADDRESS
|
||||
|
||||
CRTC = $E80000
|
||||
CRTC_R04 = CRTC+4*2 ; V total, in scanlines, minus one
|
||||
CRTC_R20 = CRTC+20*2 ; mode; bit 4 = 31.5 kHz
|
||||
HFREQ = 31500 ; lines/s in the 31.5 kHz modes. Exact: the
|
||||
; 768-wide IPL mode is 34.776 MHz / 1104 dots
|
||||
; and the 256-wide mode 11.592 MHz / 368, both
|
||||
; 31500.0 (tools/bench/crtc_mode.lua).
|
||||
|
||||
; --- state. PACE is stream.s's, deliberately: the whole point is that the
|
||||
; 68000 now writes the word the host used to write, and the pace gate that
|
||||
; reads it does not change by a single byte.
|
||||
CLK_PACE = $18034 ; == stream.s PACE
|
||||
CLK_ACC = $18060 ; word: Bresenham remainder, < HFREQ
|
||||
CLK_INCR = $18062 ; word: fps * (R04+1), computed by clk_init
|
||||
CLK_VDISP = $18064 ; long: V-DISP edges taken. An INSTRUMENT --
|
||||
; it is what lets a rig check that the tick
|
||||
; count and the raster count are the same clock.
|
||||
CLK_FPS = $18068 ; long: requested fps, an argument to clk_init
|
||||
CLK_ERR = $1806C ; long: 0 ok / 1 not a 31.5 kHz mode
|
||||
; / 2 fps*VTOTAL would overflow 16 bits
|
||||
|
||||
; ---------------------------------------------------------------------------
|
||||
; clk_init -- arm the frame clock. Reads CLK_FPS, leaves CLK_ERR.
|
||||
; Clobbers d0-d2. Leaves the CPU at SR=$2500 on success.
|
||||
; ---------------------------------------------------------------------------
|
||||
clk_init:
|
||||
clr.l CLK_ERR.l
|
||||
clr.l CLK_VDISP.l
|
||||
clr.w CLK_ACC.l
|
||||
clr.l CLK_PACE.l
|
||||
|
||||
; The mode has to be the one HFREQ describes. A 15 kHz mode would halve the
|
||||
; line rate and the divider would run at double speed while looking correct,
|
||||
; which is the failure this test exists to prevent.
|
||||
move.w CRTC_R20.l,d0
|
||||
btst #4,d0
|
||||
bne.s .modeok
|
||||
move.l #1,CLK_ERR.l
|
||||
rts
|
||||
.modeok:
|
||||
; VTOTAL and the increment. Both out of the CRTC, so a change of mode changes
|
||||
; the clock with it. acc is 16-bit and reaches at most HFREQ-1+incr, so incr
|
||||
; must leave room: 65536 - 31500 = 34036. At VTOTAL=568 that is fps < 59.9,
|
||||
; which is every rate this machine can display anyway -- but it is checked
|
||||
; rather than argued.
|
||||
move.w CRTC_R04.l,d0
|
||||
addq.w #1,d0 ; VTOTAL scanlines
|
||||
move.w d0,d1
|
||||
move.w CLK_FPS+2.l,d2 ; low word of the long
|
||||
mulu d2,d1 ; fps * VTOTAL (see FINDINGS 53.4 on
|
||||
; C68K's flat MULU charge; this is boot
|
||||
; code and is not cost-measured there)
|
||||
cmp.l #65536-HFREQ,d1
|
||||
bcs.s .fitok
|
||||
move.l #2,CLK_ERR.l
|
||||
rts
|
||||
.fitok:
|
||||
move.w d1,CLK_INCR.l
|
||||
|
||||
; The vector, before the source is armed.
|
||||
move.l #clk_isr,CLK_VEC.w
|
||||
|
||||
; Disarm everything the IPL left running, then arm GPIP4 alone. Order matters:
|
||||
; IER first (which clears IPR with it), then the edge, then the mask.
|
||||
move.b #0,MFP_IERA.l
|
||||
move.b #0,MFP_IERB.l
|
||||
move.b #0,MFP_IMRA.l
|
||||
move.b #MFP_VBASE,MFP_VR.l ; S clear: acknowledge clears pending
|
||||
bclr #MFP_GPIP4,MFP_DDR.l ; V-DISP is an input
|
||||
bclr #MFP_GPIP4,MFP_AER.l ; interrupt on the FALLING edge, i.e.
|
||||
; at the start of vertical blanking
|
||||
move.b #MFP_IVDISP,MFP_IERB.l
|
||||
move.b #MFP_IVDISP,MFP_IMRB.l
|
||||
move.w #$2500,sr ; let level 6 in; 1-5 stay masked
|
||||
rts
|
||||
|
||||
; ---------------------------------------------------------------------------
|
||||
; clk_stop -- disarm, and put the mask back where the rigs expect it.
|
||||
; ---------------------------------------------------------------------------
|
||||
clk_stop:
|
||||
move.w #$2700,sr
|
||||
move.b #0,MFP_IERB.l
|
||||
move.b #0,MFP_IMRB.l
|
||||
rts
|
||||
|
||||
; ---------------------------------------------------------------------------
|
||||
; clk_isr -- one V-DISP. Every instruction here is charged to every frame the
|
||||
; decoder draws, so it is deliberately the shortest thing that is still exact:
|
||||
; four word operations and one long increment.
|
||||
;
|
||||
; Only the LOW WORD of d0 is touched, so only the low word is saved. The
|
||||
; accumulator, the increment and the threshold all fit in 16 bits by the check
|
||||
; in clk_init, which is what makes that legal.
|
||||
; ---------------------------------------------------------------------------
|
||||
clk_isr:
|
||||
move.w d0,-(sp)
|
||||
addq.l #1,CLK_VDISP.l
|
||||
move.w CLK_ACC.l,d0
|
||||
add.w CLK_INCR.l,d0
|
||||
cmp.w #HFREQ,d0
|
||||
bcs.s .nf
|
||||
sub.w #HFREQ,d0
|
||||
addq.l #1,CLK_PACE.l
|
||||
.nf:
|
||||
move.w d0,CLK_ACC.l
|
||||
move.w (sp)+,d0
|
||||
rte
|
||||
@@ -0,0 +1,59 @@
|
||||
; Front-end for the frame clock (ROADMAP P3), for the rig.
|
||||
;
|
||||
; It exists to answer two questions that the streaming rig cannot answer on its
|
||||
; own, because there the clock is buried under a decoder:
|
||||
;
|
||||
; 1. does the tick actually come from the raster, and at exactly the rate
|
||||
; asked for -- measured over thousands of refreshes, not four;
|
||||
; 2. WHAT IT COSTS, in clocks, per interrupt. This project's currency is
|
||||
; 68000 clocks and the decoder already occupies 86.7% of the bus, so a
|
||||
; frame clock is not free until someone has priced it.
|
||||
;
|
||||
; THE INSTRUMENT, and why it is a busy loop. MAME's Lua only sees the machine
|
||||
; at frame boundaries, so it can time to 1/55.46 s and no finer -- 18 ms, where
|
||||
; the whole per-frame cost of this clock is microseconds. Differencing two
|
||||
; wall timings would measure nothing. So the 68000 counts instead: a loop with
|
||||
; ONE instruction in its body runs for a fixed number of refreshes, and the
|
||||
; iteration count is read out at both ends.
|
||||
;
|
||||
; clock off: iters0 * L = clocks in the window -> L
|
||||
; clock on: iters1 * L + ints * H = clocks in the window -> H
|
||||
;
|
||||
; The window is an exact number of raster frames, so its length in clocks is
|
||||
; exact and does not depend on the host at all. L is calibrated out by the
|
||||
; first run rather than assumed from a cycle table, which matters: the point of
|
||||
; the exercise is to price this code on the machine that will run it, and a
|
||||
; table is the thing being checked. With ~2.7e7 iterations behind it, L carries
|
||||
; enough digits that the interrupt cost -- 0.08% of the window -- survives the
|
||||
; subtraction.
|
||||
;
|
||||
; The body is `addq.l #1,CGCNT.l` and nothing else: no compare, no counter in a
|
||||
; register that an interrupt could be accused of disturbing, and a value that
|
||||
; the host can read at any moment without stopping the CPU.
|
||||
;
|
||||
; The gate does NOT decode anything. What the clock does to a real frame is
|
||||
; tools/bench/stream.lua's question, with DLX_PACE=2.
|
||||
|
||||
CGFLAG = $18070 ; 0 idle / 1 running / $EE clk_init refused
|
||||
CGON = $18074 ; 1 = arm the frame clock, 0 = leave it off
|
||||
CGCNT = $18078 ; <- loop iterations, read by the host at both
|
||||
; ends of the window
|
||||
|
||||
org $10000
|
||||
start:
|
||||
clr.l CGCNT.l
|
||||
move.l CGON.l,d0
|
||||
beq.s noclk
|
||||
bsr clk_init
|
||||
tst.l CLK_ERR.l
|
||||
bne.s bad
|
||||
noclk:
|
||||
move.l #1,CGFLAG.l ; the host starts its window here
|
||||
loop:
|
||||
addq.l #1,CGCNT.l
|
||||
bra.s loop
|
||||
bad:
|
||||
move.l #$EE,CGFLAG.l
|
||||
hold: bra.s hold
|
||||
|
||||
include "src/player/clock.i"
|
||||
@@ -4,7 +4,8 @@
|
||||
; does nothing the shipping player would not do, so that the bytes being
|
||||
; measured are the bytes that will ship. The player's own boot path will call
|
||||
; do_load once with the mode bits set to 3; this repeats it LITER times so a
|
||||
; host clock with 1/55.46 s granularity can time a job that takes milliseconds,
|
||||
; host clock with 1/56.69 s granularity (tools/bench/crtc_mode.lua) can time a
|
||||
; job that takes milliseconds,
|
||||
; and splits it by LMODE so the codebook expansion and the palette pack can be
|
||||
; priced apart. A player calls do_load with mode 7 once at boot -- the three
|
||||
; scratch tables describe the machine, not the scene -- and with mode 3 at every
|
||||
|
||||
@@ -81,6 +81,18 @@ PACE = $18034 ; producer -> decoder: frame ticks elapsed since
|
||||
; release. Frame i may not START before tick i.
|
||||
PACEON = $18038 ; 1 = obey PACE. 0 leaves the loop free-running,
|
||||
; byte for byte the loop FINDINGS 49 measured.
|
||||
LATEFR = $18080 ; frames that reached the pace gate with their
|
||||
; tick ALREADY ARRIVED, i.e. did not idle for a
|
||||
; single poll -- the previous frame used its
|
||||
; whole slot. See the gate below.
|
||||
LATEMAX = $18084 ; the worst of those, in WHOLE ticks overrun
|
||||
LATE1ST = $18088 ; index of the FIRST such frame, so that a
|
||||
; count can be told apart from a start-up
|
||||
; transient without re-running anything
|
||||
CLKON = $1803C ; 1 = the 68000 paces ITSELF: src/player/clock.i
|
||||
; drives PACE off the CRTC's V-DISP instead of
|
||||
; the host writing it. Needs PACEON=1; the gate
|
||||
; below cannot tell the two apart and must not.
|
||||
DESC = $18100 ; DESCN x u32, record base addresses
|
||||
|
||||
DESCN = 64 ; power of two; the index is masked, not compared
|
||||
@@ -92,12 +104,27 @@ SPINMAX = 2000000 ; polls with no progress before giving up
|
||||
|
||||
org $10000
|
||||
start:
|
||||
; ---- the frame clock, if this run is asking the 68000 to keep its own time.
|
||||
; It goes here rather than inside the frame loop because clk_init CLEARS PACE:
|
||||
; tick 0 has to be the instant the decoder was released, exactly as it is when
|
||||
; the host writes PACE, or the first frame's deadline moves.
|
||||
tst.l CLKON.l
|
||||
beq.s noclk
|
||||
bsr clk_init
|
||||
tst.l CLK_ERR.l
|
||||
beq.s noclk
|
||||
move.l #$E2,FLAG.l ; the clock refused; CLK_ERR says why
|
||||
bra hold
|
||||
noclk:
|
||||
move.l #1,FLAG.l ; timer starts here
|
||||
outer:
|
||||
move.l NFR.l,SCR_N.l
|
||||
clr.l FR_TAIL.l
|
||||
clr.l STALLS.l
|
||||
clr.l SPINS.l
|
||||
clr.l LATEFR.l
|
||||
clr.l LATEMAX.l
|
||||
move.l #-1,LATE1ST.l
|
||||
frameloop:
|
||||
; ---- PACE GATE (FINDINGS 49.7.2, and it is the whole point of this session).
|
||||
; Free-running, this loop asks for record i the instant it finishes record i-1,
|
||||
@@ -117,8 +144,44 @@ frameloop:
|
||||
;
|
||||
; It also makes STALLS mean something. Free-running, a stall is EARLINESS
|
||||
; (49.6); paced, a frame that has to wait for its record is a real underrun.
|
||||
;
|
||||
; AND IT COUNTS THE FRAMES THAT WERE ALREADY LATE, which is a question only a
|
||||
; REAL frame clock raises. 12 fps on a 55.4577 Hz raster is 4.6215 refreshes
|
||||
; per frame, so the divider hands out slots of 4 refreshes (72.13 ms) and 5
|
||||
; (90.16 ms), 37.9% of them short -- and the SHORT one is 13.4% under the
|
||||
; 83.33 ms every budget in this project is priced against (FINDINGS 54). A
|
||||
; frame that does not fit its slot does not fail here: PACE has already moved
|
||||
; on, so the next frame starts the instant this one finishes and the clock
|
||||
; catches up by itself. What it costs is one late PRESENT, and nothing in this
|
||||
; tree counted those because until now the tick was a host model with no
|
||||
; cadence in it at all.
|
||||
;
|
||||
; The test is "did this frame have to WAIT", not "is it a whole tick behind".
|
||||
; Frame i waits while PACE < i; so PACE >= i on arrival means the decoder came
|
||||
; to the gate with slot i already open and idled for zero polls, which is the
|
||||
; same statement as "frame i-1 ran to the end of its slot". A whole tick of
|
||||
; overrun -- PACE - FR_TAIL >= 1 -- is the much rarer case where it ran past
|
||||
; the end of the NEXT one, and is reported separately as the worst seen.
|
||||
;
|
||||
; Frame 0 is excluded: it starts at tick 0 by definition and has no predecessor
|
||||
; to have overrun. The wait loop below is untouched -- all of this is ahead of
|
||||
; it, and the free-running path executes none of it.
|
||||
tst.l PACEON.l
|
||||
beq.s nopace
|
||||
move.l PACE.l,d0
|
||||
cmp.l FR_TAIL.l,d0 ; PACE < FR_TAIL: the slot has not come
|
||||
bcs.s pacewait ; round yet, so this frame is EARLY
|
||||
tst.l FR_TAIL.l
|
||||
beq.s pacewait ; frame 0 starts AT tick 0 by definition
|
||||
tst.l LATEFR.l
|
||||
bne.s .nf1
|
||||
move.l FR_TAIL.l,LATE1ST.l
|
||||
.nf1:
|
||||
addq.l #1,LATEFR.l
|
||||
sub.l FR_TAIL.l,d0 ; whole ticks overrun; 0 = inside the
|
||||
cmp.l LATEMAX.l,d0 ; slot but with nothing left of it
|
||||
bls.s pacewait
|
||||
move.l d0,LATEMAX.l
|
||||
pacewait:
|
||||
move.l PACE.l,d0
|
||||
cmp.l FR_TAIL.l,d0 ; d0 - FR_TAIL; carry = tick not reached
|
||||
@@ -176,9 +239,17 @@ nostall:
|
||||
bne frameloop
|
||||
subq.l #1,ITER.l
|
||||
bne outer
|
||||
; ---- leave the MFP as it was found. A rig that exits with a live interrupt
|
||||
; source and a lowered mask hands the next thing that runs an interrupt it
|
||||
; has no vector for, and the failure would land somewhere else entirely.
|
||||
tst.l CLKON.l
|
||||
beq.s noclk2
|
||||
bsr clk_stop
|
||||
noclk2:
|
||||
move.l #$FF,FLAG.l ; timer stops here
|
||||
hold: bra.s hold
|
||||
desync: move.l #$EE,FLAG.l
|
||||
bra.s hold
|
||||
|
||||
include "src/player/frame.i"
|
||||
include "src/player/clock.i"
|
||||
|
||||
@@ -0,0 +1,254 @@
|
||||
#!/usr/bin/env python3
|
||||
"""What a real frame clock can be built from, and what its cadence costs.
|
||||
|
||||
python3 tools/analysis/23_frame_clock.py [--fps 12] [--vtotal 568]
|
||||
[--csv tmp/c68k_frames.csv]
|
||||
|
||||
ROADMAP P3 says "needs MFP timer or VBL", which hides the fact that ONE OF
|
||||
THOSE CANNOT DO IT and the other cannot do it either without a divider. This
|
||||
file enumerates the space rather than asserting a conclusion, the way FINDINGS
|
||||
47 had to be re-done once a hardware "no" turned out to be a claim about a whole
|
||||
configuration space nobody had walked.
|
||||
|
||||
EVERY CONSTANT HERE IS SOURCED, and from a file on this machine:
|
||||
|
||||
MFP timer clock 16 MHz / 4 MAME 0.277 sharp/x68k.cpp:1027-1028
|
||||
prescaler ladder 4,10,16,50,64,100,200 machine/mc68901.cpp:173
|
||||
timer data reg 8 bits, 0 means 256 machine/mc68901.cpp (TCDR/TADR)
|
||||
V-DISP -> GPIP4 x68k.cpp:1139, and it is also Timer A's event input,
|
||||
mc68901.cpp:167 GPIO_TIMER = {GPIP_4, GPIP_3}
|
||||
line rate 31,500 Hz exactly in both 31.5 kHz modes, derived in
|
||||
tools/bench/crtc_mode.lua from the dot clocks
|
||||
interrupt cost MEASURED, not tabled: tools/bench/clock_run.sh
|
||||
|
||||
THE PART THAT IS NOT A CLOCK PROBLEM AT ALL. 12 fps on a 55.4577 Hz raster is
|
||||
4.6215 refreshes per frame, so every frame is shown for 4 refreshes or 5 --
|
||||
72.13 ms or 90.16 ms -- and 37.9% of them get the short one. That is the
|
||||
display's quantisation and no choice of clock changes it. What it changes is
|
||||
the BUDGET: 833,333 clocks is the mean slot, not the slot, and the short slot is
|
||||
721,270. With the per-frame decode costs in hand this file says exactly how
|
||||
many frames do not fit theirs, which is a thing this project has never had to
|
||||
ask because until now the tick came from a host that could not miss.
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
import os
|
||||
import sys
|
||||
|
||||
MFP_HZ = 16_000_000 // 4 # x68k.cpp:1027-1028
|
||||
PRESCALER = [4, 10, 16, 50, 64, 100, 200] # mc68901.cpp:173
|
||||
HFREQ = 31500 # lines/s, crtc_mode.lua
|
||||
CPUHZ = 10_000_000 # 40 MHz / 4, x68k.cpp:1133
|
||||
|
||||
|
||||
def timer_space(fps):
|
||||
"""Every (prescale, data) the MFP can be set to, against a target fps."""
|
||||
slowest = MFP_HZ / (PRESCALER[-1] * 256)
|
||||
print(f"\n=== 1. THE MFP TIMER, WHICH CANNOT DO IT ALONE")
|
||||
print(f" timer clock {MFP_HZ:,} Hz, prescalers {PRESCALER}, data 1..256")
|
||||
print(f" slowest tick any single timer can produce: "
|
||||
f"{MFP_HZ}/({PRESCALER[-1]}*256) = {slowest:.3f} Hz")
|
||||
print(f" a {fps} fps frame needs {fps} Hz, which is {slowest/fps:.1f}x "
|
||||
f"slower than that -- so a software divider is REQUIRED whatever the "
|
||||
f"source, and 'use an MFP timer' is not by itself an answer.")
|
||||
exact = [(p, d) for p in PRESCALER for d in range(1, 257)
|
||||
if (MFP_HZ * d * p) and (MFP_HZ % (p * d) == 0)
|
||||
and (MFP_HZ // (p * d)) % fps == 0]
|
||||
print(f" settings whose tick rate is a whole multiple of {fps} Hz, so that "
|
||||
f"a plain counter would be exact: {len(exact)}")
|
||||
if exact:
|
||||
best = min(exact, key=lambda pd: MFP_HZ / (pd[0] * pd[1]))
|
||||
p, d = best
|
||||
tick = MFP_HZ / (p * d)
|
||||
print(f" slowest of them: prescale /{p} data {d} = {tick:.4f} Hz, "
|
||||
f"{tick/fps:.0f} ticks per frame")
|
||||
print(f" -> {tick/fps:.0f} interrupts per frame, against 4.6215 for "
|
||||
f"the raster: {tick/fps/(HFREQ/(fps*568)):.1f}x the cost, and its "
|
||||
f"phase against the raster is arbitrary, so a frame would be "
|
||||
f"presented mid-scan.")
|
||||
else:
|
||||
print(f" NONE. {MFP_HZ}/{fps} = {MFP_HZ/fps:,.2f} is not an "
|
||||
f"integer, so no prescale/data pair divides to {fps} Hz at all.")
|
||||
|
||||
|
||||
def raster_space(fps, vtotal):
|
||||
hz = HFREQ / vtotal
|
||||
print(f"\n=== 2. THE RASTER, WHICH IS THE RIGHT SOURCE AND IS ALSO NOT "
|
||||
f"A WHOLE DIVIDE")
|
||||
print(f" V-DISP is {HFREQ}/{vtotal} = {hz:.4f} Hz, and it is BOTH the "
|
||||
f"GPIP4 interrupt and Timer A's event-count input")
|
||||
print(f" whole divides -- all the MFP can do in hardware, no software:")
|
||||
for n in (3, 4, 5, 6):
|
||||
f = hz / n
|
||||
print(f" Timer A event count = {n}: {f:7.4f} fps "
|
||||
f"({100*(f/fps-1):+6.2f}% from {fps})")
|
||||
print(f" {fps} fps needs {hz/fps:.4f} refreshes per frame, which is not a "
|
||||
f"whole number, so no event-count setting is exact either.")
|
||||
print(f"\n THE DIVIDER THAT IS EXACT: add fps*VTOTAL = {fps*vtotal} per "
|
||||
f"V-DISP, emit a tick at {HFREQ}, keep the remainder.")
|
||||
print(f" long-run rate = {fps}*{vtotal}/{vtotal} = {fps} fps EXACTLY, "
|
||||
f"with a remainder that never accumulates")
|
||||
print(f" the accumulator stays under {HFREQ + fps*vtotal:,}, so it is "
|
||||
f"16-bit arithmetic on a 68000 (the ceiling is fps < "
|
||||
f"{(65536-HFREQ)/vtotal:.1f} at this VTOTAL, and clk_init checks it)")
|
||||
|
||||
|
||||
def divider_gaps(fps, vtotal, n):
|
||||
"""The tick sequence src/player/clock.i emits, in refreshes per tick."""
|
||||
acc, gaps, since = 0, [], 0
|
||||
while len(gaps) < n:
|
||||
acc += fps * vtotal
|
||||
since += 1
|
||||
if acc >= HFREQ:
|
||||
acc -= HFREQ
|
||||
gaps.append(since)
|
||||
since = 0
|
||||
return gaps
|
||||
|
||||
|
||||
def host_gaps(fps, refresh_hz, n):
|
||||
"""The tick sequence tools/bench/stream.lua's HOST clock emits.
|
||||
|
||||
It looks uniform in the source -- `floor((t - t_rel) * fps)` -- and is not.
|
||||
Lua only sees the machine at frame boundaries, so tick k lands on the first
|
||||
refresh at or after k/fps, and the gaps between ticks come out as the same
|
||||
two whole numbers of refreshes the divider produces. The host-paced runs of
|
||||
FINDINGS 49 and 51 therefore already had this cadence in them; what ROADMAP
|
||||
P3 changes is who produces it, not whether it exists.
|
||||
"""
|
||||
at = [-(-int(k * refresh_hz * 1000000 // fps) // 1000000) for k in range(n + 1)]
|
||||
return [at[k + 1] - at[k] for k in range(n)]
|
||||
|
||||
|
||||
def cadence(fps, vtotal, costs, label, refresh_hz, gaps, uniform=False):
|
||||
"""Charge each frame the slot it really gets, and run the pace gate."""
|
||||
rpf = 1.0 if uniform else HFREQ / (fps * vtotal)
|
||||
lo, hi = (1, 1) if uniform else (int(rpf), int(rpf) + 1)
|
||||
slot_lo = lo / refresh_hz * CPUHZ
|
||||
slot_hi = hi / refresh_hz * CPUHZ
|
||||
nominal = CPUHZ / fps
|
||||
|
||||
n_lo = sum(1 for g in gaps[:len(costs)] if g == lo)
|
||||
print(f"\n --- {label}: refresh {refresh_hz:.4f} Hz")
|
||||
print(f" slots are {lo} refreshes = {slot_lo:,.0f} clk "
|
||||
f"({1000*lo/refresh_hz:.2f} ms) or {hi} = {slot_hi:,.0f} clk "
|
||||
f"({1000*hi/refresh_hz:.2f} ms)")
|
||||
print(f" the nominal {fps} fps budget every figure in this project is "
|
||||
f"priced against is {nominal:,.0f} clk; the SHORT slot is "
|
||||
f"{100*(slot_lo/nominal-1):+.1f}% of it")
|
||||
over_lo = sum(1 for c in costs if c > slot_lo)
|
||||
over_hi = sum(1 for c in costs if c > slot_hi)
|
||||
over_nom = sum(1 for c in costs if c > nominal)
|
||||
print(f" frames that do not fit: {over_lo}/{len(costs)} the short "
|
||||
f"slot, {over_hi}/{len(costs)} the long one, {over_nom}/{len(costs)} "
|
||||
f"the nominal budget")
|
||||
|
||||
# The schedule, with the catch-up the pace gate actually performs: frame i
|
||||
# starts at max(finish of i-1, tick i). A frame that overruns does not fail
|
||||
# -- it eats the next frame's idle, and the clock catches up by itself.
|
||||
t, tick, late, worst, first = 0.0, 0.0, 0, 0.0, None
|
||||
for i, c in enumerate(costs):
|
||||
if t <= tick:
|
||||
t = tick # idled: on time
|
||||
else:
|
||||
if i: # frame 0 has no predecessor
|
||||
late += 1
|
||||
if first is None:
|
||||
first = i
|
||||
worst = max(worst, t - tick)
|
||||
t += c
|
||||
tick += gaps[i] / refresh_hz * CPUHZ
|
||||
print(f" through the pace gate: {late}/{len(costs)} frames found "
|
||||
f"their slot already open (first at frame {first}), worst start "
|
||||
f"{worst:,.0f} clk = {1000*worst/CPUHZ:.1f} ms behind its tick")
|
||||
if not uniform:
|
||||
print(f" {n_lo}/{len(costs)} slots were the short one "
|
||||
f"({100*n_lo/len(costs):.1f}%; the exact share is "
|
||||
f"{100*(hi-rpf):.1f}%)")
|
||||
return late
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--fps", type=int, default=12)
|
||||
ap.add_argument("--vtotal", type=int, default=568,
|
||||
help="CRTC R04+1; 568 is the 31.5 kHz 256-line mode")
|
||||
ap.add_argument("--csv", default="tmp/c68k_frames.csv",
|
||||
help="per-frame decode cost, from tools/bench/c68k/run.sh")
|
||||
a = ap.parse_args()
|
||||
|
||||
timer_space(a.fps)
|
||||
raster_space(a.fps, a.vtotal)
|
||||
|
||||
print(f"\n=== 3. WHAT THE CADENCE COSTS")
|
||||
if not os.path.exists(a.csv):
|
||||
print(f" {a.csv} not found -- run tools/bench/c68k/run.sh first. The "
|
||||
f"cadence question CANNOT be answered from percentiles: it needs "
|
||||
f"the per-frame series, because what matters is whether an "
|
||||
f"expensive frame lands in a short slot and how long the catch-up "
|
||||
f"takes afterwards.")
|
||||
return 1
|
||||
costs = [float(r["cycles"]) for r in csv.DictReader(open(a.csv))]
|
||||
print(f" {len(costs)} frames from {a.csv}: mean {sum(costs)/len(costs):,.0f} "
|
||||
f"clk, max {max(costs):,.0f} (frame {costs.index(max(costs))})")
|
||||
|
||||
hw = HFREQ / a.vtotal
|
||||
n = len(costs)
|
||||
# MAME's screen is fast by htotal/(htotal-8): refresh_mode() builds the
|
||||
# frame period from scr.max_x*scr.max_y with scr.max_x = m_htotal - 8. Both
|
||||
# rasters are run, because the rig measures against the fast one and the
|
||||
# player will run on the other -- reporting only one would leave the rig's
|
||||
# count and this file's differing with nobody able to say which was wrong.
|
||||
htotal = 368
|
||||
mame = hw * htotal / (htotal - 8)
|
||||
|
||||
# The budget model every figure in FINDINGS assumes: a slot of exactly
|
||||
# 1/fps. No machine has this; it is the yardstick, run through the same
|
||||
# schedule so that what the cadence ADDS can be read off.
|
||||
cadence(a.fps, a.vtotal, costs, "the NOMINAL model (a slot of exactly "
|
||||
"1/fps, which no raster produces)", float(a.fps),
|
||||
[1] * (n + 1), uniform=True)
|
||||
cadence(a.fps, a.vtotal, costs, "the HOST tick, as tools/bench/stream.lua "
|
||||
"actually emits it", mame, host_gaps(a.fps, mame, n + 1))
|
||||
cadence(a.fps, a.vtotal, costs, "the 68000's own clock on the HARDWARE "
|
||||
"raster", hw, divider_gaps(a.fps, a.vtotal, n + 1))
|
||||
cadence(a.fps, a.vtotal, costs, f"the 68000's own clock on MAME's raster "
|
||||
f"(fast by {htotal}/{htotal-8})", mame,
|
||||
divider_gaps(a.fps, a.vtotal, n + 1))
|
||||
|
||||
print(f"""
|
||||
Reading it.
|
||||
|
||||
THE CADENCE WAS ALREADY THERE. The nominal row is the model every budget in
|
||||
this project is priced against -- a slot of exactly 1/fps -- and no raster
|
||||
produces it. The host row is what tools/bench/stream.lua has been emitting all
|
||||
along: `floor((t - t_rel) * fps)` looks uniform, but Lua only sees the machine at
|
||||
frame boundaries, so its ticks land on refreshes and its gaps are the same two
|
||||
whole numbers. The host-paced results of FINDINGS 49 and 51 therefore already
|
||||
carried a 4/5 cadence that nothing named. ROADMAP P3 did not introduce it; it
|
||||
moved who produces it onto the machine, where it belongs, and made it visible.
|
||||
|
||||
THE SHORT SLOT IS REAL AND IT IS NOT A FAILURE. {sum(1 for c in costs if c > 721270)}/{len(costs)} frames do not fit
|
||||
721,270 clocks. The pace gate only says "not before tick i", so a frame that
|
||||
overruns spends the next frame's idle and the clock recovers by itself; the cost
|
||||
is one frame presented a refresh late, not a dropped frame. What the counts
|
||||
above measure is frames with no idle left, and the difference between the
|
||||
nominal row and the raster rows -- 1 against 4 -- is the whole price of the
|
||||
cadence on this container.
|
||||
|
||||
THE EXPENSIVE FRAME IS FRAME 0, at {max(costs)/(CPUHZ/a.fps)*100:.0f}% of the nominal budget: the first
|
||||
frame of a scene has nothing to SKIP against, so it is the whole picture in one
|
||||
slot. Most of what follows it in these counts is that transient draining, which
|
||||
is why the first index is printed next to the total. It also means the cost is
|
||||
paid AT A SCENE CHANGE, alongside the 18.96 ms of loader (FINDINGS 53.2) and the
|
||||
seek -- not spread over the window.
|
||||
|
||||
SCOPE. Decode costs are C68K's, on zero-wait-state memory, so they are a lower
|
||||
bound; real DRAM moves every row here in the same direction. The MAME rows are
|
||||
the emulator's fast raster and exist to be compared with tools/bench/pace_run.sh,
|
||||
not to describe hardware.""")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -11,7 +11,8 @@
|
||||
-- a LOWER BOUND, not a prediction. Interrupts are masked (SR=$2700) so the
|
||||
-- IPL's timer and VBL handlers cannot steal cycles into the measurement.
|
||||
--
|
||||
-- Timing resolution is one video frame (1/55.46 s = 18.03 ms), because Lua
|
||||
-- Timing resolution is one video frame (1/56.69 s = 17.64 ms -- MAME's, not
|
||||
-- the hardware's 55.46; see crtc_mode.lua), because Lua
|
||||
-- gets no cycle counter -- luaengine.cpp exposes machine.time and nothing
|
||||
-- from device_execute_interface. Each variant therefore loops enough times
|
||||
-- to run ~4 emulated seconds, putting the granularity error near 0.4%.
|
||||
|
||||
+45
-1
@@ -286,7 +286,7 @@ echo "--- session 21: the 68000 builds its own codebooks and palette (FINDINGS 5
|
||||
#
|
||||
# NOT gated on the cycle counts, and the reason is NOT the one blit.s has. These
|
||||
# are emulated time and reproduce exactly run to run; what they are not is
|
||||
# sharp, because MAME samples them on a 1/55.46 s clock and the job takes
|
||||
# sharp, because MAME samples them on a 1/56.69 s clock and the job takes
|
||||
# milliseconds. Nothing in the tree's cost models depends on them either. A
|
||||
# change in them is a re-derivation in FINDINGS 53, not a red light here.
|
||||
bash tools/bench/load_run.sh "$DLX" > tmp/load_gate.log 2>&1 || {
|
||||
@@ -294,4 +294,48 @@ bash tools/bench/load_run.sh "$DLX" > tmp/load_gate.log 2>&1 || {
|
||||
exit 1; }
|
||||
grep -aE "^ *OK|both CPU cores|SCENE CHANGE" tmp/load_gate.log | sed 's/^ *//;s/^/ /'
|
||||
|
||||
echo "--- session 22: the 68000 keeps its own frame clock (FINDINGS 54) ---"
|
||||
# ROADMAP P3. Until now the 12 fps tick came from tools/bench/stream.lua -- a
|
||||
# host writing a word into emulated RAM. A player has no host. src/player/
|
||||
# clock.i derives the tick from the CRTC's own V-DISP output through the MFP,
|
||||
# with a remainder-keeping divider whose two constants are READ OUT OF THE CRTC
|
||||
# at init, so the clock and the raster it counts cannot disagree.
|
||||
#
|
||||
# WHAT IS GATED, and it is deliberately structural rather than numeric:
|
||||
# * the interrupt count equals the raster frame count -- the tick IS the
|
||||
# raster, not something that merely resembles it;
|
||||
# * the divider does not accumulate drift, stated in TICKS (a remainder can
|
||||
# hold back at most one) rather than in ppm, which would let a longer
|
||||
# window advertise a tighter clock for free;
|
||||
# * every frame tick waits 4 or 5 refreshes and nothing else, which is what a
|
||||
# remainder-keeping divider can produce and a broken one cannot.
|
||||
# The interrupt COST is printed and not gated, for the same reason FINDINGS 53's
|
||||
# cycle counts are not: it is a measurement, and a change in it is a
|
||||
# re-derivation in FINDINGS 54 rather than a red light here.
|
||||
bash tools/bench/clock_run.sh 3000 12 > tmp/clock_gate.log 2>&1 || {
|
||||
echo "FAIL: the frame clock did not pass."; tail -12 tmp/clock_gate.log
|
||||
exit 1; }
|
||||
grep -aE "INTERRUPT:|PER FRAME:|DRIFT:|CADENCE:" tmp/clock_gate.log
|
||||
grep -q "V-DISP interrupts 3000" tmp/clock_gate.log || {
|
||||
echo "FAIL: the tick is not the raster -- the interrupt count and the frame"
|
||||
echo " count disagree. Everything else in FINDINGS 54 rests on that."
|
||||
exit 1; }
|
||||
|
||||
echo "--- session 22: 120 frames decoded on the machine's own clock (FINDINGS 54) ---"
|
||||
# The strongest form of the claim: the same pixel-exact 120-frame decode out of
|
||||
# the same 256 KB ring, with NOTHING outside the machine deciding when a frame
|
||||
# may start. The pace gate in src/player/stream.s is byte-for-byte the one
|
||||
# FINDINGS 51 measured -- it cannot tell a host-written tick from a machine-
|
||||
# written one, which is why this is a test of the clock and not of a new rig.
|
||||
DLX_PACE=2 bash tools/bench/pace_run.sh 256 0 > tmp/selfpace_check.log 2>&1 || {
|
||||
echo "FAIL: the self-paced pass did not complete."; tail -8 tmp/selfpace_check.log
|
||||
exit 1; }
|
||||
grep -aE "decoder SELF-PACED|FRAME CLOCK|UNDERRUNS|NO IDLE" tmp/selfpace_check.log
|
||||
grep -q "UNDERRUNS: 0/120" tmp/selfpace_check.log || {
|
||||
echo "FAIL: the self-paced decoder underran."; exit 1; }
|
||||
grep -q "^OK" tmp/selfpace_check.log || {
|
||||
echo "FAIL: the self-paced pass was not pixel-exact. The clock changed WHEN"
|
||||
echo " frames start; if it changed WHAT they draw, the interrupt is"
|
||||
echo " corrupting decoder state."; tail -4 tmp/selfpace_check.log; exit 1; }
|
||||
|
||||
echo "ALL GREEN"
|
||||
|
||||
@@ -0,0 +1,240 @@
|
||||
-- Drive src/player/clockgate.s: measure the 68000's own FRAME CLOCK.
|
||||
-- ROADMAP P3.
|
||||
--
|
||||
-- Two things are being measured and they need different instruments.
|
||||
--
|
||||
-- THE RATE AND THE CADENCE are counted, not timed. The clock's tick is a
|
||||
-- V-DISP interrupt, and MAME's Lua sees the machine once per screen frame --
|
||||
-- which is once per V-DISP. So the host's sampling granularity is exactly the
|
||||
-- clock's own granularity, and the cadence comes out as integers: how many
|
||||
-- refreshes each frame tick waited. There is no timing error to argue about
|
||||
-- in a count of 4s and 5s.
|
||||
--
|
||||
-- THE COST IS TIMED BY THE 68000, because the host cannot. 1/55.46 s of host
|
||||
-- granularity is 18 ms and the interrupt costs microseconds. So the 68000 runs
|
||||
-- a one-instruction loop for a window of thousands of refreshes and the host
|
||||
-- reads the iteration count at both ends; the interrupt cost falls out of the
|
||||
-- difference between a run with the clock armed and one without. See the head
|
||||
-- of src/player/clockgate.s for the arithmetic. This script emits the raw
|
||||
-- counts; tools/bench/clock_cost.py does the subtraction, so that the two runs
|
||||
-- it needs can be separate MAME invocations.
|
||||
--
|
||||
-- MEASUREMENT SCOPE. This is MAME 0.277's emulated X68000, not real hardware.
|
||||
-- What is being priced is the interrupt sequence of MAME's cycle-accurate
|
||||
-- M68000 core (src/devices/cpu/m68000, the `M68000` device x68k.cpp:1133 asks
|
||||
-- for) against zero-wait-state RAM. Real DRAM adds wait states to the six bus
|
||||
-- cycles of the exception and the four of the handler alike, so this is a LOWER
|
||||
-- BOUND in the same way every other 68000 figure in this project is.
|
||||
--
|
||||
-- Env:
|
||||
-- DLX_CLK_ON 1 = arm the frame clock, 0 = leave it off (the calibration
|
||||
-- run). REQUIRED -- the two runs are not interchangeable and a
|
||||
-- default would let one be reported as the other.
|
||||
-- DLX_CLK_FPS frame rate to ask clk_init for (default 12)
|
||||
-- DLX_CLK_WIN measurement window, in raster frames (default 3000 = 54.1 s)
|
||||
-- DLX_CLK_OUT where to write the raw counts (default tmp/clock_run.txt)
|
||||
|
||||
M = manager.machine
|
||||
SP = M.devices[":maincpu"].spaces["program"]
|
||||
|
||||
local function findfile(n)
|
||||
for _,p in ipairs{"../tools/bench/"..n, "tools/bench/"..n, n} do
|
||||
local f = io.open(p,"rb"); if f then f:close(); return p end
|
||||
end
|
||||
error(n.." not found")
|
||||
end
|
||||
local MODE = loadfile(findfile("crtc_mode.lua"))()
|
||||
|
||||
local CGFLAG, CGON, CGCNT = 0x18070, 0x18074, 0x18078
|
||||
local CLK_PACE = 0x18034
|
||||
local CLK_ACC, CLK_INCR = 0x18060, 0x18062
|
||||
local CLK_VDISP, CLK_FPS = 0x18064, 0x18068
|
||||
local CLK_ERR = 0x1806C
|
||||
local CPUHZ = 10000000
|
||||
|
||||
local ONS = os.getenv("DLX_CLK_ON")
|
||||
local FPS = tonumber(os.getenv("DLX_CLK_FPS") or "") or 12
|
||||
local WIN = tonumber(os.getenv("DLX_CLK_WIN") or "") or 3000
|
||||
local OUT = os.getenv("DLX_CLK_OUT") or "clock_run.txt"
|
||||
|
||||
local function P(s) print("[CLK] "..s) end
|
||||
|
||||
if ONS ~= "0" and ONS ~= "1" then
|
||||
P("DLX_CLK_ON must be 0 (calibration, clock off) or 1 (clock armed). The "
|
||||
.."cost figure is the DIFFERENCE between the two runs, so neither is "
|
||||
.."meaningful alone and neither gets to be the default.")
|
||||
M:exit()
|
||||
return
|
||||
end
|
||||
local ON = (ONS == "1")
|
||||
|
||||
local code do local f=io.open("clockgate.bin","rb"); code=f:read("a"); f:close() end
|
||||
|
||||
local function T() local t=M.time; return t.seconds + t.attoseconds/1e18 end
|
||||
|
||||
-- Settling frames between the gate reporting `running` and the window opening.
|
||||
-- The CPU may still be inside clk_init when the host first sees CGFLAG=1, and
|
||||
-- the first V-DISP edge after arming lands wherever the raster happens to be.
|
||||
-- Two frames puts the window entirely inside the steady state.
|
||||
local SETTLE = 2
|
||||
|
||||
local st, n = "boot", 0
|
||||
local f_ready, f0, f1 = nil, nil, nil
|
||||
local c0, c1, v0, v1, p0, p1, t0, t1
|
||||
-- Cadence: refreshes between consecutive frame ticks. Recorded as a histogram
|
||||
-- and as the raw first few, because the interesting claim is not the mean (the
|
||||
-- divider makes that exact by construction) but that the SPREAD is only ever
|
||||
-- the two values either side of fps*VTOTAL/HFREQ.
|
||||
local last_pace, last_pace_f, cad, seen_tick = nil, nil, {}, false
|
||||
|
||||
SUB = emu.add_machine_frame_notifier(function()
|
||||
local ok, err = pcall(function()
|
||||
n = n + 1
|
||||
if st == "boot" then
|
||||
if T() < 3.0 then return end
|
||||
MODE.apply(SP)
|
||||
for i = 1, #code do SP:write_u8(0x10000+i-1, string.byte(code,i)) end
|
||||
SP:write_u32(CGFLAG, 0)
|
||||
SP:write_u32(CGON, ON and 1 or 0)
|
||||
SP:write_u32(CLK_FPS, FPS)
|
||||
local cpu = M.devices[":maincpu"]
|
||||
cpu.state["SR"].value = 0x2700 -- clk_init lowers it to $2500 itself
|
||||
cpu.state["SP"].value = 0x8000
|
||||
cpu.state["PC"].value = 0x10000
|
||||
P(string.format("clockgate.bin=%d B, clock %s, asking for %d fps, "
|
||||
.."window %d raster frames", #code,
|
||||
ON and "ARMED" or "OFF (calibration run)", FPS, WIN))
|
||||
st = "wait"; return
|
||||
end
|
||||
if st == "wait" then
|
||||
local fl = SP:read_u32(CGFLAG)
|
||||
if fl == 0xEE then
|
||||
local e = SP:read_u32(CLK_ERR)
|
||||
P("clk_init REFUSED: CLK_ERR="..e..(e == 1 and
|
||||
" (CRTC is not in a 31.5 kHz mode, so HFREQ=31500 would be wrong)" or
|
||||
e == 2 and " (fps*VTOTAL does not fit the 16-bit accumulator)" or ""))
|
||||
M:exit(); return
|
||||
end
|
||||
if fl ~= 1 then
|
||||
if T() > 60 then P("TIMEOUT: the gate never started"); M:exit() end
|
||||
return
|
||||
end
|
||||
f_ready = n; st = "settle"; return
|
||||
end
|
||||
if st == "settle" then
|
||||
if n < f_ready + SETTLE then return end
|
||||
f0, t0 = n, T()
|
||||
c0 = SP:read_u32(CGCNT)
|
||||
v0 = SP:read_u32(CLK_VDISP)
|
||||
p0 = SP:read_u32(CLK_PACE)
|
||||
last_pace, last_pace_f = p0, n
|
||||
if ON then
|
||||
P(string.format("armed: incr=%d (fps*VTOTAL), acc=%d, first tick "
|
||||
.."pending", SP:read_u16(CLK_INCR),
|
||||
SP:read_u16(CLK_ACC)))
|
||||
end
|
||||
st = "run"; return
|
||||
end
|
||||
if st == "run" then
|
||||
if ON then
|
||||
local pc = SP:read_u32(CLK_PACE)
|
||||
if pc ~= last_pace then
|
||||
-- The FIRST change is dropped. Its interval runs from the window
|
||||
-- opening rather than from a tick, so it measures where the window
|
||||
-- happened to start and would show up as a spurious short bucket.
|
||||
if seen_tick then
|
||||
-- More than one tick in a single refresh would mean fps above the
|
||||
-- raster rate; give it its own bucket rather than averaging it in.
|
||||
local gap = n - last_pace_f
|
||||
if pc - last_pace > 1 then gap = 0 end
|
||||
cad[gap] = (cad[gap] or 0) + 1
|
||||
end
|
||||
seen_tick = true
|
||||
last_pace, last_pace_f = pc, n
|
||||
end
|
||||
end
|
||||
if n < f0 + WIN then return end
|
||||
f1, t1 = n, T()
|
||||
c1 = SP:read_u32(CGCNT)
|
||||
v1 = SP:read_u32(CLK_VDISP)
|
||||
p1 = SP:read_u32(CLK_PACE)
|
||||
st = "done"
|
||||
|
||||
local frames = f1 - f0
|
||||
local secs = t1 - t0
|
||||
local clocks = secs * CPUHZ
|
||||
local iters = c1 - c0
|
||||
local ints = v1 - v0
|
||||
local ticks = p1 - p0
|
||||
P(string.format("window: %d raster frames, %.6f s emulated -> %.0f "
|
||||
.."68000 clocks", frames, secs, clocks))
|
||||
-- THE INSTRUMENT IS 2.22% FAST AND IT IS WORTH SAYING SO EVERY RUN.
|
||||
-- The CRTC registers describe a 31,500 lines/s raster of VTOTAL lines.
|
||||
-- MAME does not run it at that rate: x68k_crtc.cpp refresh_mode()
|
||||
-- computes the frame period as (scr.max_x * scr.max_y) dots with
|
||||
-- scr.max_x = m_htotal - 8, one character cell short and an INCLUSIVE
|
||||
-- rectangle bound used as a count. So the emulated raster is fast by
|
||||
-- htotal/(htotal-8) -- 368/360 in this mode -- and every rate derived
|
||||
-- from it here is fast by the same factor. The divider under test is
|
||||
-- built on the registers, so its HARDWARE rate is the asked-for one and
|
||||
-- what this rig can check is that it tracks whatever raster it is given.
|
||||
local vtotal = SP:read_u16(0xE80008) + 1
|
||||
local htotal = (SP:read_u16(0xE80000) + 1) * 8
|
||||
local hw_hz = 31500 / vtotal
|
||||
local skew = htotal / (htotal - 8)
|
||||
P(string.format(" raster period %.4f ms = %.4f Hz", 1000*secs/frames,
|
||||
frames/secs))
|
||||
P(string.format(" the CRTC registers describe 31500/%d = %.4f Hz; "
|
||||
.."MAME is fast by htotal/(htotal-8) = %d/%d = %.4f",
|
||||
vtotal, hw_hz, htotal, htotal-8, skew))
|
||||
P(string.format(" loop iterations %d", iters))
|
||||
if ON then
|
||||
P(string.format(" V-DISP interrupts %d, frame ticks %d", ints,
|
||||
ticks))
|
||||
-- The self-check that makes the rest of it worth reading: the interrupt
|
||||
-- count and the host's screen-frame count are supposed to be the SAME
|
||||
-- clock seen from two sides. If they disagree by more than the one
|
||||
-- edge the window boundaries can straddle, the tick is not the raster.
|
||||
if math.abs(ints - frames) > 1 then
|
||||
P(string.format("FAIL: %d V-DISP interrupts over %d raster frames. "
|
||||
.."The tick is not coming from the raster.", ints,
|
||||
frames))
|
||||
M:exit(); return
|
||||
end
|
||||
-- Two numbers, and confusing them is the whole trap. The measured rate
|
||||
-- is against MAME's fast raster; dividing the skew out gives the rate
|
||||
-- the same code produces on a machine whose raster matches its own
|
||||
-- registers, which is the number the player is judged on.
|
||||
local meas = ticks/secs
|
||||
P(string.format(" measured rate %.6f fps against MAME's raster "
|
||||
.."(%+.0f ppm vs the asked %d)", meas,
|
||||
1e6*(meas/FPS - 1), FPS))
|
||||
P(string.format(" de-skewed %.6f fps -> %+.1f ppm from %d, "
|
||||
.."which is the tick quantisation of %d ticks and not "
|
||||
.."drift", meas/skew, 1e6*(meas/skew/FPS - 1), FPS,
|
||||
ticks))
|
||||
local ks = {}
|
||||
for k in pairs(cad) do ks[#ks+1] = k end
|
||||
table.sort(ks)
|
||||
local s = ""
|
||||
for _,k in ipairs(ks) do
|
||||
s = s .. string.format("%d:%d ", k, cad[k])
|
||||
end
|
||||
P(" cadence, refreshes per frame tick: "..s)
|
||||
end
|
||||
|
||||
local fh = io.open(OUT, "w")
|
||||
fh:write(string.format("on %d\nfps %d\nframes %d\nsecs %.15g\n"
|
||||
.."clocks %.15g\niters %d\nints %d\nticks %d\n"
|
||||
.."vtotal %d\nhtotal %d\nhw_hz %.15g\nskew %.15g\n",
|
||||
ON and 1 or 0, FPS, frames, secs, clocks, iters,
|
||||
ints, ticks, vtotal, htotal, hw_hz, skew))
|
||||
for k, v in pairs(cad) do fh:write(string.format("cad %d %d\n", k, v)) end
|
||||
fh:close()
|
||||
P("counts -> "..OUT)
|
||||
P("done")
|
||||
M:exit(); return
|
||||
end
|
||||
end)
|
||||
if not ok then print("[CLK] LUA ERROR: "..tostring(err)); M:exit() end
|
||||
end)
|
||||
@@ -0,0 +1,132 @@
|
||||
"""What the 68000's own frame clock costs, out of the two clock.lua runs.
|
||||
|
||||
ROADMAP P3. Usage: clock_cost.py <off-run.txt> <on-run.txt>
|
||||
|
||||
THE SUBTRACTION. Both runs execute the same one-instruction loop over a window
|
||||
of the same number of raster frames, so the window is the same number of 68000
|
||||
clocks in both. With the clock off, every clock in the window went into loop
|
||||
iterations:
|
||||
|
||||
L = clocks / iters_off clocks per iteration
|
||||
|
||||
With it armed, the interrupts took some of them:
|
||||
|
||||
H = (clocks - iters_on * L) / ints clocks per V-DISP interrupt
|
||||
|
||||
L is CALIBRATED rather than looked up. That is the point: this project's cost
|
||||
model (tools/analysis/buscost.py) says a 68000 bus cycle is 4 clocks and an
|
||||
instruction costs 4 * (instruction words + data accesses), and the whole reason
|
||||
to measure is to avoid scoring the clock against the table the table is meant to
|
||||
be checked by. L falling on a whole number of clocks is therefore a RESULT, not
|
||||
an assumption, and it is reported as one.
|
||||
|
||||
WHAT THE FIGURE IS PER FRAME. Not H -- the interrupt fires once per refresh and
|
||||
a frame is several refreshes. On the hardware raster that is 31500/VTOTAL over
|
||||
fps interrupts per frame, and the de-skewed rate is the one to use: MAME's
|
||||
raster is fast by htotal/(htotal-8) (see tools/bench/clock.lua), and charging
|
||||
the player the emulator's extra interrupts would overstate the cost by that
|
||||
same 2.2%.
|
||||
"""
|
||||
import sys
|
||||
|
||||
|
||||
def read(path):
|
||||
d, cad = {}, {}
|
||||
for line in open(path):
|
||||
f = line.split()
|
||||
if f[0] == "cad":
|
||||
cad[int(f[1])] = int(f[2])
|
||||
else:
|
||||
d[f[0]] = float(f[1])
|
||||
d["cad"] = cad
|
||||
return d
|
||||
|
||||
|
||||
def main(off_path, on_path):
|
||||
off, on = read(off_path), read(on_path)
|
||||
if off["on"] != 0 or on["on"] != 1:
|
||||
sys.exit("FAIL: expected the calibration run first and the armed run "
|
||||
"second; got on=%d then on=%d" % (off["on"], on["on"]))
|
||||
for k in ("frames", "clocks", "fps", "vtotal"):
|
||||
if off[k] != on[k]:
|
||||
sys.exit("FAIL: the two runs do not share a window: %s is %g in "
|
||||
"the calibration run and %g in the armed one"
|
||||
% (k, off[k], on[k]))
|
||||
|
||||
clocks = off["clocks"]
|
||||
L = clocks / off["iters"]
|
||||
ints = on["ints"]
|
||||
H = (clocks - on["iters"] * L) / ints
|
||||
|
||||
# The self-check that licenses the subtraction: the interrupt count must be
|
||||
# the raster frame count. clock.lua already fails on this, restated here
|
||||
# because this file is also read on its own.
|
||||
if abs(ints - on["frames"]) > 1:
|
||||
sys.exit("FAIL: %d interrupts over %g raster frames -- not the raster"
|
||||
% (ints, on["frames"]))
|
||||
|
||||
fps, skew = on["fps"], on["skew"]
|
||||
hw_hz = on["hw_hz"]
|
||||
per_frame_ints = hw_hz / fps
|
||||
per_frame = H * per_frame_ints
|
||||
FRAME_CLK = 10e6 / fps
|
||||
|
||||
print(" calibration: %.6f clocks per loop iteration over %d iterations"
|
||||
% (L, off["iters"]))
|
||||
print(" (%s a whole number of clocks -- the loop is one "
|
||||
"`addq.l #1,abs.l` at 7 bus cycles plus a `bra.s`)"
|
||||
% ("lands on" if abs(L - round(L)) < 1e-3 else "does NOT land on"))
|
||||
print(" INTERRUPT: %.2f clocks per V-DISP, measured over %d of them"
|
||||
% (H, ints))
|
||||
print(" PER FRAME: %.2f interrupts x %.2f = %.0f clocks = %.4f%% of a "
|
||||
"%g fps frame" % (per_frame_ints, H, per_frame,
|
||||
100 * per_frame / FRAME_CLK, fps))
|
||||
print(" (%.4f refreshes per frame on the HARDWARE raster of "
|
||||
"31500/%d = %.4f Hz, not on MAME's, which is %.4fx fast)"
|
||||
% (per_frame_ints, on["vtotal"], hw_hz, skew))
|
||||
|
||||
# THE DRIFT GATE, and it is stated in TICKS rather than in ppm on purpose.
|
||||
# A remainder-keeping divider emits floor() or ceil() of the exact tick
|
||||
# count over any window and never accumulates -- so the only honest
|
||||
# tolerance is one tick, and any ppm figure is that one tick divided by
|
||||
# however long the window happened to be. Quoting ppm would let a longer
|
||||
# window advertise a tighter clock for no reason.
|
||||
want = on["frames"] * fps * on["vtotal"] / 31500.0
|
||||
ticks = on["ticks"]
|
||||
print(" DRIFT: %d ticks over %d refreshes; exact is %.4f, so the "
|
||||
"error is %+.4f ticks" % (ticks, on["frames"], want, ticks - want))
|
||||
if abs(ticks - want) > 1.0:
|
||||
sys.exit("FAIL: %d ticks where %.4f were due -- off by %.2f, which is "
|
||||
"more than the one tick a remainder can hold back. The "
|
||||
"divider is accumulating drift." % (ticks, want, ticks - want))
|
||||
|
||||
cad = on["cad"]
|
||||
tot = sum(cad.values())
|
||||
if tot:
|
||||
# Refreshes per frame is 31500 / (fps * VTOTAL) exactly -- the divider's
|
||||
# own ratio, upside down. A remainder-keeping divider can only ever
|
||||
# emit the two whole numbers either side of it, so anything else in the
|
||||
# histogram is a bug in the divider and not a rounding taste.
|
||||
rpf = 31500.0 / (fps * on["vtotal"])
|
||||
lo, hi = int(rpf), int(rpf) + 1
|
||||
print(" CADENCE: %s (%d intervals; %.4f refreshes per frame, so "
|
||||
"only %d and %d are possible)"
|
||||
% (", ".join("%dx%d (%.1f%%)" % (k, v, 100.0 * v / tot)
|
||||
for k, v in sorted(cad.items())), tot, rpf, lo, hi))
|
||||
for k in cad:
|
||||
if k not in (lo, hi):
|
||||
sys.exit("FAIL: a frame tick waited %d refreshes, which a "
|
||||
"remainder-keeping divider cannot produce" % k)
|
||||
# The mix is forced too: lo*a + hi*b = refreshes, a + b = ticks.
|
||||
b = tot * rpf - lo * tot
|
||||
print(" expected %d:%d split %.1f%% / %.1f%%, got "
|
||||
"%.1f%% / %.1f%%"
|
||||
% (lo, hi, 100 * (tot - b) / tot, 100 * b / tot,
|
||||
100.0 * cad.get(lo, 0) / tot, 100.0 * cad.get(hi, 0) / tot))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) != 3:
|
||||
sys.exit(__doc__)
|
||||
sys.exit(main(sys.argv[1], sys.argv[2]))
|
||||
Executable
+36
@@ -0,0 +1,36 @@
|
||||
#!/bin/bash
|
||||
# One frame-clock run: the 68000 derives its own 12 fps tick from the raster
|
||||
# (ROADMAP P3, FINDINGS 54).
|
||||
#
|
||||
# tools/bench/clock_run.sh [window-in-raster-frames] [fps]
|
||||
#
|
||||
# TWO MAME INVOCATIONS, and they are not interchangeable. The first leaves the
|
||||
# clock off and calibrates the cost of the gate's own loop; the second arms it.
|
||||
# The interrupt cost is the difference, so a run that reported only the second
|
||||
# would be reporting a number it cannot compute. See src/player/clockgate.s.
|
||||
#
|
||||
# Only MAME can run this: the frame clock is an MFP interrupt driven by the
|
||||
# CRTC's V-DISP output, and tools/bench/c68k has neither device. That is why
|
||||
# this stage has no second-core half, unlike load_run.sh.
|
||||
set -e
|
||||
cd "$(dirname "$0")/../.."
|
||||
WIN=${1:-3000}
|
||||
FPS=${2:-12}
|
||||
|
||||
tools/vasm/vasmm68k_mot -Fbin -o tmp/clockgate.bin src/player/clockgate.s > /dev/null
|
||||
|
||||
for ON in 0 1; do
|
||||
# stdbuf -oL: without it a long MAME run is unobservable until it exits, and
|
||||
# a run that is merely finishing looks exactly like one that is wedged (34.1).
|
||||
( cd tmp && DLX_CLK_ON=$ON DLX_CLK_FPS=$FPS DLX_CLK_WIN=$WIN \
|
||||
DLX_CLK_OUT=clock_$ON.txt SDL_VIDEODRIVER=dummy stdbuf -oL \
|
||||
timeout -k 5 600 mame x68000 -bios ipl10 -ramsize 2M -video soft -window \
|
||||
-sound none -nothrottle -plugins -autoboot_script ../tools/bench/clock.lua \
|
||||
-seconds_to_run 240 > clock_$ON.log 2>&1 )
|
||||
# A run that never reached the counts must fail as that, not as bad arithmetic.
|
||||
grep -q "^\[CLK\] done" tmp/clock_$ON.log || {
|
||||
echo "FAIL: the clock rig did not finish run ON=$ON -- no completion marker."
|
||||
tail -8 tmp/clock_$ON.log; exit 1; }
|
||||
done
|
||||
grep -a "^\[CLK\]" tmp/clock_1.log | sed -n '/window:/,/cadence/p' | sed 's/\[CLK\] / /'
|
||||
python3 tools/bench/clock_cost.py tmp/clock_0.txt tmp/clock_1.txt
|
||||
@@ -23,6 +23,26 @@
|
||||
-- Total blanking time is identical to the 768 mode (112 dots @ 11.592MHz =
|
||||
-- 336 dots @ 34.776MHz = 9.66us), which is what a real monitor needs.
|
||||
--
|
||||
-- THE EMULATOR DOES NOT RUN THE RASTER THESE REGISTERS DESCRIBE, and every
|
||||
-- rig in this tree samples the machine at ITS rate, not at the hardware's.
|
||||
-- x68k_crtc.cpp refresh_mode() builds the frame period as
|
||||
--
|
||||
-- (scr.max_x * scr.max_y) dots / dotclock, scr.max_x = m_htotal - 8
|
||||
--
|
||||
-- which is one character cell short AND uses an inclusive rectangle bound as a
|
||||
-- count. So MAME's refresh is fast by htotal/(htotal-8) = 368/360 = 1.02222:
|
||||
-- 56.6901 Hz where the registers say 55.4577. MEASURED, not read off the
|
||||
-- source alone -- tools/bench/clock.lua reports both every run, and they agree
|
||||
-- to six digits (FINDINGS 54.5).
|
||||
--
|
||||
-- It matters in exactly two places and is harmless in the rest. Anything timed
|
||||
-- by counting host frames has 1/56.69 s of granularity, not 1/55.46; and
|
||||
-- anything PACED by the raster runs 2.22% fast under MAME. It does NOT touch
|
||||
-- 68000 cycle figures: the CPU clock is 40 MHz/4 and has nothing to do with the
|
||||
-- screen. Do not "correct" the 55.4577 below to match a measurement -- it is
|
||||
-- the hardware's, derived from the dot clocks above, and it is what
|
||||
-- src/player/clock.i builds its divider on.
|
||||
--
|
||||
-- VERTICAL registers are NOT halved. The CRTC still generates a 568-line
|
||||
-- 31.5kHz raster (31500/568 = 55.46 Hz); "256 lines" is a graphics-layer
|
||||
-- double-scan (draw_gfx() halves gfxrect, x68k_v.cpp:401). Halving them would
|
||||
|
||||
@@ -93,7 +93,7 @@ end
|
||||
|
||||
-- The plan: one sequential correctness pass, then the cost anchors, then a
|
||||
-- full pass timed. Iteration counts target ~4 emulated seconds each so the
|
||||
-- 1/55.46 s timing granularity costs under 0.5%.
|
||||
-- 1/56.69 s timing granularity (crtc_mode.lua) costs under 0.5%.
|
||||
-- DLX_VERIFY_ONLY=1 drops the cost anchors and runs only the correctness pass,
|
||||
-- so tools/bench/check.sh can gate the decoder without paying for ~2 minutes of
|
||||
-- timing runs that would make the green light sensitive to host load anyway.
|
||||
|
||||
@@ -18,7 +18,7 @@
|
||||
--
|
||||
-- MEASUREMENT SCOPE, unchanged from decode.lua: MAME's memory carries no wait
|
||||
-- states, so these are pure 68000 instruction cycles -- a LOWER BOUND on real
|
||||
-- hardware. Interrupts are masked (SR=$2700). The host clock has 1/55.46 s
|
||||
-- hardware. Interrupts are masked (SR=$2700). The host clock has 1/56.69 s
|
||||
-- granularity and the job takes milliseconds, so each configuration is repeated
|
||||
-- LITER times and divided; repeating is honest because do_load is not
|
||||
-- temporally recursive -- every pass rewrites what the last one wrote, from the
|
||||
|
||||
+13
-2
@@ -15,12 +15,23 @@
|
||||
#
|
||||
# kbps 0 = unlimited pipe. There is no default rate anywhere in this tree
|
||||
# (FINDINGS 50) and there is none here either.
|
||||
#
|
||||
# DLX_PACE selects WHO KEEPS THE TIME: 1 (default) is the host writing the tick,
|
||||
# 2 is the 68000 writing it off the CRTC's V-DISP (ROADMAP P3, FINDINGS 54).
|
||||
# Everything else about the run is identical, which is the whole point -- the
|
||||
# pace gate in src/player/stream.s cannot tell them apart, so a difference in
|
||||
# the result is a difference in the CLOCK and not in the rig.
|
||||
set -e
|
||||
cd "$(dirname "$0")/../.."
|
||||
RING=${1:?ring KB}; KBPS=${2:?pipe KB/s, or 0 for unlimited}
|
||||
CUT_AT=$3; CUT_FR=${4:-1}
|
||||
DLX=${DLX:-tmp/rc_fr_singe_scsi_span.dlx}
|
||||
PACE=${DLX_PACE:-1}
|
||||
TAG="r${RING}_k${KBPS}${CUT_AT:+_cut${CUT_AT}x${CUT_FR}}"
|
||||
# The default tag is left ALONE when the host keeps the time: tools/bench/
|
||||
# pace_sweep.sh reads tmp/pace_r<ring>_k<kbps>.log by name, and renaming the
|
||||
# host-paced logs would break a sweep that has nothing to do with this option.
|
||||
if [ "$PACE" != 1 ]; then TAG="${TAG}_p$PACE"; fi
|
||||
|
||||
tools/vasm/vasmm68k_mot -Fbin -o tmp/stream.bin src/player/stream.s > /dev/null
|
||||
[ -f tmp/stream_disk.bin ] || python3 tools/bench/prep_stream.py "$DLX" > tmp/prep_stream.log
|
||||
@@ -29,7 +40,7 @@ mkdir -p "tmp/snap_pace_$TAG"; rm -f "tmp/snap_pace_$TAG/x68000"/*.png
|
||||
# of a prefix is not an assignment token, so bash takes the next word as the
|
||||
# command and the run dies with "SDL_VIDEODRIVER=dummy: command not found".
|
||||
CUTENV=(); [ -n "$CUT_AT" ] && CUTENV=(DLX_CUT_AT="$CUT_AT" DLX_CUT_FR="$CUT_FR")
|
||||
( cd tmp && env DLX_PACE=1 DLX_RING_KB=$RING DLX_STREAM_KBPS=$KBPS \
|
||||
( cd tmp && env DLX_PACE=$PACE DLX_RING_KB=$RING DLX_STREAM_KBPS=$KBPS \
|
||||
"${CUTENV[@]}" DLX_SLACK_CSV="slack_$TAG.csv" \
|
||||
SDL_VIDEODRIVER=dummy stdbuf -oL timeout -k 5 900 \
|
||||
mame x68000 -bios ipl10 -ramsize 2M -video soft -window -sound none \
|
||||
@@ -42,6 +53,6 @@ grep -q "snapshot taken" "tmp/pace_$TAG.log" || {
|
||||
echo "FAIL($TAG): no snapshot marker -- the pass did not complete."
|
||||
tail -6 "tmp/pace_$TAG.log"; exit 1; }
|
||||
echo "=== $TAG"
|
||||
grep -aE "decoder (PACED|FREE)|ring: |UNDERRUNS|SEEK SLACK|RING-BOUND|RATE-BOUND|BUILD TIME|PIPE CUT|DEADLINE|REQUIRED" \
|
||||
grep -aE "decoder (SELF-PACED|PACED|FREE)|FRAME CLOCK|ring: |UNDERRUNS|NO IDLE|SEEK SLACK|RING-BOUND|RATE-BOUND|BUILD TIME|PIPE CUT|DEADLINE|REQUIRED" \
|
||||
"tmp/pace_$TAG.log" | sed "s/\[STR\] / /"
|
||||
python3 tools/bench/verify_decode.py "$DLX" --snap "tmp/snap_pace_$TAG" | tail -2
|
||||
|
||||
@@ -107,7 +107,8 @@ end
|
||||
local function launch(cfg)
|
||||
push(STREAM, blob, cfg.off+1, cfg.len)
|
||||
clear_picture()
|
||||
-- ~4 emulated seconds per config: 1/55.46 s granularity costs under 0.5%.
|
||||
-- ~4 emulated seconds per config: 1/56.69 s granularity (crtc_mode.lua)
|
||||
-- costs under 0.5%.
|
||||
local est = cfg.nspans*(cfg.var == 5 and 60 or (cfg.var == 6 and 50 or 70))
|
||||
+ cfg.npix*10
|
||||
cfg.iter = math.max(4, math.floor(4*CPUHZ/est))
|
||||
|
||||
+84
-7
@@ -53,7 +53,17 @@
|
||||
-- (FINDINGS 42.1). A default is how a folklore number ends up
|
||||
-- silently underneath a table nobody restates it in.
|
||||
-- DLX_PREFILL_KB bytes to deliver before releasing the CPU (default 0)
|
||||
-- DLX_PACE 1 = hold the decoder to META.fps (default 0 = free-run)
|
||||
-- DLX_PACE 1 = hold the decoder to META.fps, with the HOST writing
|
||||
-- the tick (default 0 = free-run)
|
||||
-- 2 = hold it to META.fps with the 68000 writing its own
|
||||
-- tick, off the CRTC's V-DISP (src/player/clock.i, ROADMAP
|
||||
-- P3). Same gate, same ring, same deadlines; what changes
|
||||
-- is that nothing outside the machine decides when a frame
|
||||
-- may start. Deadlines are then taken from the ticks the
|
||||
-- machine actually emitted rather than from a host model of
|
||||
-- 12 fps -- which matters, because MAME's raster runs 2.22%
|
||||
-- fast (see tools/bench/clock.lua) and a host model would
|
||||
-- quietly grade every arrival against the wrong clock.
|
||||
-- DLX_CUT_AT frame tick at which the pipe stops dead (a seek). Needs
|
||||
-- DLX_PACE; unset = no cut.
|
||||
-- DLX_CUT_FR how many frame times the cut lasts (default 1)
|
||||
@@ -88,7 +98,9 @@ local META = loadfile("stream_meta.lua")()
|
||||
local FLAG, ITER, NFR = 0x18000, 0x18008, 0x1800C
|
||||
local RD_PTR, FR_HEAD, FR_TAIL = 0x18020, 0x18024, 0x18028
|
||||
local STALLS, SPINS, DESC = 0x1802C, 0x18030, 0x18100
|
||||
local PACE, PACEON = 0x18034, 0x18038
|
||||
local PACE, PACEON, CLKON = 0x18034, 0x18038, 0x1803C
|
||||
local CLK_VDISP, CLK_FPS, CLK_ERR = 0x18064, 0x18068, 0x1806C
|
||||
local LATEFR, LATEMAX, LATE1ST = 0x18080, 0x18084, 0x18088
|
||||
local DESCN = 64
|
||||
local CB1, CB4 = 0x20000, 0x22000
|
||||
local RING = 0x40000
|
||||
@@ -108,7 +120,8 @@ if KBPS == nil then
|
||||
return
|
||||
end
|
||||
local PREFILL = (tonumber(os.getenv("DLX_PREFILL_KB") or "") or 0) * 1024
|
||||
local PACED = (os.getenv("DLX_PACE") == "1")
|
||||
local SELFCLK = (os.getenv("DLX_PACE") == "2")
|
||||
local PACED = (os.getenv("DLX_PACE") == "1") or SELFCLK
|
||||
local CUT_AT = tonumber(os.getenv("DLX_CUT_AT") or "")
|
||||
local CUT_FR = tonumber(os.getenv("DLX_CUT_FR") or "") or 1
|
||||
local SLACK_CSV = os.getenv("DLX_SLACK_CSV")
|
||||
@@ -269,9 +282,13 @@ local function setup()
|
||||
SP:write_u32(FLAG, 0); SP:write_u32(FR_HEAD, 0); SP:write_u32(FR_TAIL, 0)
|
||||
SP:write_u32(ITER, 1); SP:write_u32(NFR, META.nframes)
|
||||
SP:write_u32(PACE, 0); SP:write_u32(PACEON, PACED and 1 or 0)
|
||||
SP:write_u32(CLKON, SELFCLK and 1 or 0)
|
||||
SP:write_u32(CLK_FPS, META.fps)
|
||||
P(string.format("stream.bin=%d B, codebooks %d+%d B, disk %d B, %d frames",
|
||||
#code, META.cb1_len, META.cb4_len, META.disk_len, META.nframes))
|
||||
P(string.format("decoder %s%s", PACED and ("PACED at "..META.fps.." fps")
|
||||
P(string.format("decoder %s%s", SELFCLK
|
||||
and ("SELF-PACED at "..META.fps.." fps off V-DISP")
|
||||
or PACED and ("PACED at "..META.fps.." fps by the host")
|
||||
or "FREE-RUNNING (tests wrap, not buffering -- 49.7.2)",
|
||||
CUT_AT and string.format(", pipe cut at tick %d for %.2f fr",
|
||||
CUT_AT, CUT_FR) or ""))
|
||||
@@ -292,6 +309,12 @@ local function launch()
|
||||
end
|
||||
|
||||
local st, t0, t_rel = "boot", nil, 0
|
||||
-- tick_t[i+1] = emulated time of frame tick i, filled in as they are observed.
|
||||
-- Under a self-clock this replaces the host's `t_rel + i/fps` as the deadline:
|
||||
-- the machine's clock is the one the player is actually held to, and MAME's
|
||||
-- raster is 2.22% fast, so grading arrivals against a host model of 12 fps
|
||||
-- would flatter every record by that much.
|
||||
local tick_t = {}
|
||||
local pace, min_ahead, min_at = -1, math.huge, -1
|
||||
local sum_ahead, n_ahead = 0, 0
|
||||
local slack_series = {}
|
||||
@@ -314,10 +337,21 @@ SUB = emu.add_machine_frame_notifier(function()
|
||||
end
|
||||
if st == "running" then
|
||||
if PACED then
|
||||
local tick = math.floor((t - t_rel) * META.fps)
|
||||
-- WHO WRITES THE TICK. Host-paced, the tick is a host model of
|
||||
-- META.fps and this script advances it. Self-paced, the 68000 has
|
||||
-- already advanced it off the raster and this script only READS it --
|
||||
-- the frame gate in src/player/stream.s cannot tell the two apart,
|
||||
-- which is the point: the same bytes are gated either way.
|
||||
local tick = SELFCLK and SP:read_u32(PACE)
|
||||
or math.floor((t - t_rel) * META.fps)
|
||||
if tick > pace then
|
||||
pace = tick
|
||||
SP:write_u32(PACE, pace)
|
||||
if not SELFCLK then SP:write_u32(PACE, pace) end
|
||||
-- The emulated time this tick actually happened, which is what a
|
||||
-- record's deadline is measured against under a self-clock. Ticks
|
||||
-- can arrive more than one apart if the host misses a frame, so the
|
||||
-- whole run is filled rather than just the newest.
|
||||
for k = #tick_t, tick do tick_t[k+1] = t end
|
||||
-- Taken BEFORE this tick's frame is decoded, so snapshot n is the
|
||||
-- finished picture of frame n-1. tick 0 is skipped: nothing has been
|
||||
-- drawn yet and it would record a black screen as a decoded frame.
|
||||
@@ -358,6 +392,14 @@ SUB = emu.add_machine_frame_notifier(function()
|
||||
.."or described wrongly, not that the codec changed.")
|
||||
M:exit(); return
|
||||
end
|
||||
if fl == 0xE2 then
|
||||
local e = SP:read_u32(CLK_ERR)
|
||||
P("FRAME CLOCK REFUSED: CLK_ERR="..e..(e == 1 and
|
||||
" (the CRTC is not in a 31.5 kHz mode, so the divider's 31500 "
|
||||
.."lines/s would be wrong)" or e == 2 and
|
||||
" (fps*VTOTAL does not fit the 16-bit accumulator)" or ""))
|
||||
M:exit(); return
|
||||
end
|
||||
if fl == 0xE1 then
|
||||
P("PRODUCER STALLED OUT -- stream.s spun SPINMAX times with no new "
|
||||
.."record. Delivered "..nsent.."/"..META.nframes..".")
|
||||
@@ -382,6 +424,25 @@ SUB = emu.add_machine_frame_notifier(function()
|
||||
META.nframes, dt, dt*CPUHZ/META.nframes,
|
||||
100*(dt*CPUHZ/META.nframes)/FRAME12, META.fps))
|
||||
end
|
||||
if SELFCLK then
|
||||
-- The clock's own report, in the run where it actually paced a
|
||||
-- decoder rather than a busy loop. The rate is against MAME's fast
|
||||
-- raster; tools/bench/clock_run.sh is where it gets de-skewed.
|
||||
local vd = SP:read_u32(CLK_VDISP)
|
||||
local vt = SP:read_u16(0xE80008) + 1 -- VTOTAL, as clk_init read it
|
||||
-- DELIBERATELY NOT A RATE. 120 ticks is far too short a window to
|
||||
-- quote fps from: the run's start and end each straddle a tick, so
|
||||
-- +/-1 on 119 intervals is +/-8000 ppm and would read as drift the
|
||||
-- clock does not have. What IS exact here is a count of refreshes
|
||||
-- against a count of ticks. The rate figure comes from
|
||||
-- tools/bench/clock_run.sh, over thousands of them.
|
||||
P(string.format("FRAME CLOCK: %d V-DISP interrupts drove %d ticks "
|
||||
.."(%.4f refreshes/frame; the divider's own ratio is "
|
||||
.."31500/(%d*%d) = %.4f), %.0f clocks/frame of "
|
||||
.."interrupt at FINDINGS 54's 181.35 each",
|
||||
vd, pace + 1, vd/(pace + 1), META.fps, vt,
|
||||
31500/(META.fps*vt), 181.35*vd/(pace+1)))
|
||||
end
|
||||
P(string.format("ring: %d wraps, %d B of hole (mean %.1f KB, %.1f%% of "
|
||||
.."the ring)", holes, hole_bytes,
|
||||
holes > 0 and hole_bytes/holes/1024 or 0,
|
||||
@@ -395,6 +456,20 @@ SUB = emu.add_machine_frame_notifier(function()
|
||||
-- opposite thing, so the two are never printed in the same words.
|
||||
P(string.format("UNDERRUNS: %d/%d frames waited past their %d fps slot "
|
||||
.."(%d polls)", stalls, META.nframes, META.fps, spins))
|
||||
-- The other way a paced frame can go wrong, and it is not the same
|
||||
-- failure. An UNDERRUN is the pipe: the record was not there. This
|
||||
-- is the CPU: the record was there, the slot was already open, and
|
||||
-- the decoder had no idle left in the frame before. Under a
|
||||
-- host-written tick every slot is 83.33 ms; under the machine's own
|
||||
-- clock they alternate 72.13 and 90.16 ms, 37.9% of them short, and
|
||||
-- this is what the short ones cost (FINDINGS 54).
|
||||
local late, latemax = SP:read_u32(LATEFR), SP:read_u32(LATEMAX)
|
||||
local late1 = SP:read_u32(LATE1ST)
|
||||
P(string.format("NO IDLE: %d/%d frames found their slot already open "
|
||||
.."-- the frame before used all of it; worst overrun "
|
||||
.."%d whole tick%s, first at frame %d", late,
|
||||
META.nframes, latemax, latemax == 1 and "" or "s",
|
||||
late1 == 0xFFFFFFFF and -1 or late1))
|
||||
-- WHAT THE MINIMUM OVER A RUN IS, AND IS NOT. At release the ring
|
||||
-- holds only what the prefill put there, so the early ticks report a
|
||||
-- buffer that has not been built yet, not a ring or a pipe that
|
||||
@@ -455,7 +530,9 @@ SUB = emu.add_machine_frame_notifier(function()
|
||||
for i = 0, META.nframes-1 do
|
||||
local a = arrival[i+1]
|
||||
if a then
|
||||
local late = a - (t_rel + i / META.fps)
|
||||
local due = SELFCLK and tick_t[i+1] or (t_rel + i / META.fps)
|
||||
if not due then due = t_rel + i / META.fps end
|
||||
local late = a - due
|
||||
if late > 0 then
|
||||
misses = misses + 1
|
||||
if late > worst then worst = late end
|
||||
|
||||
Reference in New Issue
Block a user