src/player/decode.s now paints v7 literal spans, pixel-exact under MAME and px68k's C68K core over a container where every frame carries 128-216 spans covering up to 38% of the picture. The span pass is blit.s v7 verbatim: the 66.0/9.143/9.978 fit was measured on that instruction sequence. The container is DLX3 -- a span section between the mode header and the block payload, since that is the only place the 68000 can reach without first parsing something of variable length. 16_span_roundtrip.py gates it in check.sh, and asserts it emitted enough spans to have tested anything. Two synthetic all-SPAN anchors price v7 inside decode.s at 151.2 and 225.6 clocks per 4x4 block, against FINDINGS 40's table of 151 and 226 -- 0.2% on both emulators. The measured mode costs what it was said to cost. Two things that were not on the list: TWO BYTE BUDGETS. FINDINGS 40's 18/120 was scored against the 488 KB/s PIPE, not the 280 KB/s profile, and at the profile rate the lam search has already spent the allowance -- spans fired on 5 frames of 120 and looked like a regression. The profile is a chosen quality rate point; the pipe is hardware. --kbps and --span-kbps are now separate and spans run before mu, because a span pays in bytes and mu pays in picture. Delivered: 86/120 over budget without spans, 77/120 at the profile budget, 34/120 on the pipe for +0.36 dB. C_SKIP_MIXED WAS NEVER MEASURED, and it was 18% low -- 45.0, now 55.0. It is the one constant in the table that came from a derivation, because the synthetic frame that would measure it cannot exist: a byte needs a coded block for its SKIP to be mixed. Four bracketing anchors measure it on both emulators with the header byte rotated through all four positions, and the partner mode solves back to its own anchored value to 0.2%. With it corrected the model predicts a real spanned decode to -0.06% mean / 0.09% worst, against -2.99% / 4.30%. It matters because a span marks its run SKIP, so mixed SKIPs dominate exactly the frames spans are judged on. Also: the rig had been writing its synthetic timing frames 26 KB past the top of a 2 MB machine, and got away with it because the modes it overran are data-independent. A span's jump displacements come out of the stream, so it is not. And frames-over-budget is no longer a safe headline -- the controller aims at the deadline, so 55 of 120 frames sit within 5% of it and a 1% cost shift moves 22 frames. FINDINGS 41. check.sh ALL GREEN, now gating on a span-heavy DLX3 container. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
191 lines
8.6 KiB
Python
191 lines
8.6 KiB
Python
"""Bus-cycle cost of src/player/decode.s and of tools/bench/blit.s's v6 spans.
|
|
|
|
A 68000 bus cycle is 4 clocks (S0-S7) with no wait states, and the 68000
|
|
prefetches every instruction word over the same bus. So a block's bus cost is
|
|
`instruction words + data accesses`, a long access counting twice on the 16-bit
|
|
bus and `movem.l` of N registers counting 2N.
|
|
|
|
The per-path word counts are read off tools/bench/decode.lst and
|
|
tools/bench/blit.s. tools/analysis/15_bus_occupancy.py checks the DATA half of
|
|
this table against tools/bench/c68k/c68k_bench, which counts every bus callback
|
|
the C68K core makes: they agree to 0.04%. The prefetch half cannot be measured
|
|
from either emulator -- MAME does not expose a fetch count and C68K reads
|
|
opcodes through a host pointer with no callback -- so it rests on that check.
|
|
"""
|
|
BUS_CLK = 4
|
|
|
|
# --- decode.s, per block ---------------------------------------------------
|
|
# dispatch move.b (a1),d0 / lsr.b / and.w #3 / beq .sk 6w, 1 read
|
|
# + subq / beq .v1 -> 8w
|
|
# + subq / bne .rw -> 10w
|
|
# V4 body $10090..$100E2 = 82 B = 41w; 4 x (1 byte read
|
|
# + movem.l 2 = 4 reads + 2 move.l = 4 writes) = 36
|
|
# V1 body $100E2..$10106 = 36 B = 18w; 1 byte read
|
|
# + movem.l 8 = 16 reads + 4 x movem.l 2 = 16 wr = 33
|
|
# RAW body $10106..$10164 = 94 B = 47w; 8 x (2 byte reads
|
|
# + 1 move.l = 2 writes) = 32
|
|
BODY = {0: (0, 0), 1: (18, 33), 2: (41, 36), 3: (47, 32)}
|
|
DISPATCH = {0: 6, 1: 8, 2: 10, 3: 10}
|
|
SK_TAIL = 1 # addq.l #8,a4
|
|
GROUP_HEAD = 3 # tst.b (a1) + beq allskip
|
|
GROUP_TAIL = 4 # addq.l #1,a1 / cmpa.l a5,a4 / bne byteloop
|
|
ALLSKIP = 9 # the whole four-block fast path, tst.b included
|
|
ROW_HEAD, ROW_TAIL = 3, 7
|
|
|
|
# --- blit.s v6 spans -------------------------------------------------------
|
|
# One chain unit moves 12 registers = 48 B = 24 pixels:
|
|
# movem.l (a0)+,12 = 2w instr + 24 word reads = 26
|
|
# movem.l 12,(a2) = 2w instr + 24 word writes = 26
|
|
# lea 48(a2),a2 = 2w instr = 2
|
|
# Per span: move.l (a0)+,a2 (1w + 2 reads) + move.w (a0)+,d0 (1w + 1 read)
|
|
# + jmp v6ch(pc,d0.w) (2w) + dbra (2w) = 9
|
|
V6_UNIT_PX = 24
|
|
V6_UNIT_BUS = 54
|
|
V6_SPAN_BUS = 9
|
|
V6_SPAN_CYC = 43.7 # MEASURED, FINDINGS 30
|
|
V6_PX_CYC = 9.152 # MEASURED, FINDINGS 30
|
|
|
|
# --- a DMAC array-chaining span -------------------------------------------
|
|
# SOURCED, MC68450 Direct Memory Access Controller, Motorola, Jul 1989
|
|
# (bitsavers). These replace session-10's first pass, which guessed 2 bus
|
|
# cycles a pixel from bus arithmetic and was 12% optimistic.
|
|
#
|
|
# Fig 4-25 sheet 4, DUAL ADDRESS / OPERAND SIZE IS WORD / DEVICE SIZE IS
|
|
# 16-BITS, D->M or M->D: {WORD READ, WORD WRITE} = 9 CLOCKS.
|
|
# Confirmed by the long-operand row: two of each = 18 clocks.
|
|
# Fig 4-25 note 2: reads are 4 clocks and WRITES ARE 5. That extra clock on
|
|
# every write is the whole story -- it is why the DMAC does not beat a 68000
|
|
# movem chain, which writes in 4.
|
|
DMA_PX_CLK = 9
|
|
# Fig 4-25 sheet 1, SEQUENTIAL ARRAY CHAINING: 36 CLOCKS per entry (three
|
|
# word reads to fetch the 6-byte entry, plus reload).
|
|
DMA_CHAIN_CLK = 36
|
|
# Sect 4.5.2.1 front-end overhead 5 clocks best case, 8 worst; 4.5.2.2
|
|
# back-end 2 clocks best. Once per period of bus ownership, not per span.
|
|
DMA_FRONT_CLK, DMA_BACK_CLK = 5, 2
|
|
# Fig 4-25 sheet 3, SINGLE ADDRESS: W/B READ 4 clocks, W/B WRITE 5 clocks.
|
|
# A device->memory disk transfer is one memory WRITE = 5 clocks if the DMAC
|
|
# holds the bus, or 5 + front + back = 12 if it arbitrates per word.
|
|
# FINDINGS 5's long-standing 8 clk/word ESTIMATE sits inside that range.
|
|
DMA_DISK_CLK_WORD_HELD, DMA_DISK_CLK_WORD_ARB = 5, 12
|
|
|
|
# The 68000 cannot execute while another master owns the bus: no cache, and a
|
|
# two-word prefetch queue that empties immediately. So DMA time is ADDITIVE to
|
|
# CPU time, not overlapped -- which is what FINDINGS 35's flat debit assumed
|
|
# and session 10's first pass wrongly "refined".
|
|
DMA_OVERLAPS = False
|
|
|
|
|
|
def pad24(npix):
|
|
return -(-npix // V6_UNIT_PX) * V6_UNIT_PX
|
|
|
|
|
|
def block_bus(mode_map, spanned=None):
|
|
"""(instruction words, data accesses) for one frame's CPU block decode.
|
|
|
|
`spanned` is a boolean array the same shape as mode_map marking blocks a
|
|
span will paint instead; those blocks still cost their dispatch, because
|
|
the mode map is walked either way, but not their body."""
|
|
nby, nbx = mode_map.shape
|
|
pref = nby * (ROW_HEAD + ROW_TAIL)
|
|
data = 0
|
|
for by in range(nby):
|
|
row = mode_map[by]
|
|
sp = spanned[by] if spanned is not None else None
|
|
for gi in range(0, nbx, 4):
|
|
g = row[gi:gi + 4]
|
|
if (g == 0).all():
|
|
pref += ALLSKIP
|
|
data += 1
|
|
continue
|
|
pref += GROUP_HEAD + GROUP_TAIL - 1 # BLOCK 0 has no lsr.b
|
|
data += 1
|
|
for k, b in enumerate(g):
|
|
b = int(b)
|
|
if sp is not None and sp[gi + k]:
|
|
b = 0 # the span paints it
|
|
pw, pd = BODY[b]
|
|
pref += DISPATCH[b] + pw + SK_TAIL
|
|
data += 1 + pd
|
|
return pref, data
|
|
|
|
|
|
# --- v7: v6 with a finer tail (MEASURED, session 11, FINDINGS 40) ----------
|
|
# v6 pads every span up to 24 pixels because its unrolled chain is built from
|
|
# 12-register movem units, and FINDINGS 39.3 attributed 86% of the DMAC array
|
|
# chain's advantage over v6 to exactly that padding. v7 keeps the coarse chain
|
|
# and appends a second chain whose unit is one `move.l (a0)+,(a2)+` -- 2 pixels,
|
|
# so the quantum is 2 and a run of 4x4 blocks pads to NOTHING.
|
|
#
|
|
# Session 10 proposed a 2-REGISTER MOVEM tail (4 pixels, derived at 56 clocks)
|
|
# and that would have been the wrong instruction: movem.l (a0)+,d0-d1 plus
|
|
# movem.l d0-d1,(a2) plus the lea is 14 bus cycles for 4 pixels, where two plain
|
|
# move.l are 10. The plainest instruction on the machine wins the tail.
|
|
#
|
|
# The second entry point needs a second dispatch, and the fine displacement is
|
|
# carried MID-STREAM (after the coarse pixels, before the fine ones) rather than
|
|
# in the span record, so the decoder holds nothing extra across the copy and
|
|
# keeps all 12 payload registers. Costed as 2 more bytes per span.
|
|
#
|
|
# MEASURED by tools/bench/span.sh (blit.s v7, 13 span lengths, every config
|
|
# pixel-exact): cycles = 66.0/span + 9.143/coarse pixel + 9.978/fine pixel,
|
|
# fitting all 13 to within 0.2%.
|
|
V7_SPAN_CYC = 66.0 # MEASURED, FINDINGS 40
|
|
V7_CPX_CYC = 9.143 # MEASURED, FINDINGS 40 (24-pixel coarse unit)
|
|
V7_FPX_CYC = 9.978 # MEASURED, FINDINGS 40 (2-pixel fine unit)
|
|
V7_FINE_PX = 2
|
|
# Bus: per span v6's 9 plus a second {move.w (a0)+,d0 ; jmp} = 2 + 2.
|
|
# Per fine unit: move.l (a0)+,(a2)+ = 1 instruction word + 2 reads + 2 writes.
|
|
V7_SPAN_BUS = 13
|
|
V7_FINE_BUS = 5
|
|
V7_SPAN_HDR = 8 # {u32 address, u16 coarse disp} + u16 fine disp
|
|
|
|
|
|
def pad2(npix):
|
|
return -(-npix // V7_FINE_PX) * V7_FINE_PX
|
|
|
|
|
|
def v7_span(npix):
|
|
"""(pixels carried, CPU clocks) for a v7 span of npix pixels."""
|
|
k, r = divmod(pad2(npix), V6_UNIT_PX)
|
|
return (k * V6_UNIT_PX + r,
|
|
V7_SPAN_CYC + k * V6_UNIT_PX * V7_CPX_CYC + r * V7_FPX_CYC)
|
|
|
|
|
|
def v7_span_bus(npix):
|
|
"""Bus CYCLES a v7 span occupies -- instruction words plus data accesses."""
|
|
p, d = v7_span_split(npix)
|
|
return p + d
|
|
|
|
|
|
def v7_span_split(npix):
|
|
"""(instruction words, data accesses) for one v7 span, separately.
|
|
|
|
15_bus_occupancy.py needs the two apart, because the DATA half is what the
|
|
C68K harness can check and the PREFETCH half is what rides on that check.
|
|
|
|
per span move.l (a0)+,a2 1 word + 2 reads
|
|
move.w (a0)+,d0 1 word + 1 read (coarse displacement)
|
|
jmp (pc,d0.w) 2 words
|
|
move.w (a0)+,d0 1 word + 1 read (fine, from mid-stream)
|
|
jmp (pc,d0.w) 2 words
|
|
dbra 2 words -> 9 words, 4 accesses
|
|
per coarse 2 movem.l of 12 + lea = 6 words, 24 reads + 24 writes
|
|
per fine move.l (a0)+,(a2)+ = 1 word, 2 reads + 2 writes
|
|
"""
|
|
k, r = divmod(pad2(npix), V6_UNIT_PX)
|
|
f = r // V7_FINE_PX
|
|
return (9 + k * 6 + f * 1,
|
|
4 + k * 48 + f * 4)
|
|
|
|
|
|
# Per FRAME, decode.s's paint_spans entry and exit: the span count read, the
|
|
# guard branch, and the push/pop of a1 that buys back a twelfth payload
|
|
# register. Two long accesses a frame against 24 pixels a chain unit.
|
|
V7_FRAME_PREF, V7_FRAME_DATA = 7, 7
|
|
|
|
|
|
def v6_span_bus(npix):
|
|
k = pad24(npix) // V6_UNIT_PX
|
|
return V6_SPAN_BUS + k * V6_UNIT_BUS
|