First 68000 instructions in this project to draw a pixel. Everything before this was GVRAM filled from Lua, which costs zero 68000 cycles, so the blit figure the whole CPU budget rests on had never been validated. Four variants of a full-frame 256x192 paint, timed in MAME and each also hand-derived from the MC68000 timing tables beforehand; the two agree to 0.006-0.43%, which is what makes the result trustworthy after this project's history of false-good measurements. V1 movem.l blit from a word-expanded RAM frame 446,286 cyc 53.6% V2 naive move.b/move.w per pixel 1,284,174 cyc 154.1% V3 write-only floor, no source read 225,789 cyc 27.1% V4 same writes in 4x4 block order 637,971 cyc 76.6% Scope: MAME's gvram_w/gvram_r carry no timing at all, so these are instruction cycles against zero-wait-state memory -- a floor, not a hardware prediction. V1's output snapshots pixel-exact through verify_frame256.py, closing FINDINGS 23.5. The V1/V3 gap shows reading the source frame is exactly half the cost, which makes the architecture question live: decode-direct-to-GVRAM needs no RAM reference frame and scales with the non-SKIP block fraction, crossing compose-then-blit at 70% of blocks changed. That fraction is now the top priority and is already a by-product of vq_hybrid.py's mode decision. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
157 lines
6.3 KiB
ArmAsm
157 lines
6.3 KiB
ArmAsm
; Full-frame GVRAM blit cost on a stock 68000 @ 10MHz.
|
|
;
|
|
; Answers: what fraction of a 12fps frame budget (833,333 cycles) does simply
|
|
; PUTTING a decoded 256x192 frame on screen cost, before any decoding?
|
|
;
|
|
; Geometry (tools/bench/crtc_mode.lua): 256-colour page, one pixel per WORD of
|
|
; CPU address space, 1024-byte line stride, picture in rows 32..223 of a
|
|
; 256-row page. So a row is 512 contiguous bytes of writes, then a 512-byte
|
|
; skip. 192 rows = 98,304 bytes of GVRAM write traffic per frame.
|
|
;
|
|
; Confirmed from MAME 0.277 x68k_crtc.cpp:501 (gvram_w, case 0x0100): a CPU
|
|
; write in 256-colour mode is masked to 0x00ff, so the HIGH byte of every word
|
|
; written is discarded by the hardware. V1 exploits this -- it never has to
|
|
; clear the odd bytes of its source.
|
|
;
|
|
; Three variants, selected by VAR, each looped ITER times:
|
|
; V1 movem.l blit from a word-expanded RAM frame (96KB). The realistic
|
|
; "decode to RAM, then blit" design. Reads 96KB, writes 96KB.
|
|
; V2 naive byte-source expansion (move.b / move.w per pixel). The obvious
|
|
; implementation, kept as the baseline V1 has to beat.
|
|
; V3 write-only floor: registers preloaded once, no source read at all.
|
|
; Nothing that puts this many pixels on screen can beat V3. The gap
|
|
; V1-V3 is the price of reading a source frame at all.
|
|
; V4 the SAME 96KB of writes, but issued in 4x4 BLOCK order instead of
|
|
; row-linear order. This is the access pattern a decoder that writes
|
|
; codewords straight into GVRAM actually has, and it is the number that
|
|
; picks the decoder architecture: compose-in-RAM-then-blit (V1) versus
|
|
; decode-direct-to-GVRAM (V4 scaled by the fraction of non-SKIP blocks).
|
|
; Each block is 4 rows of 8 bytes at a 1024-byte stride, so the
|
|
; destination displacements 0/1024/2048/3072 all fit a 16-bit offset and
|
|
; the block needs only one base pointer. V4 deliberately scrambles the
|
|
; picture (it reads a row-linear source in block order); it is a timing
|
|
; probe, which is why the correctness snapshot is taken after V1.
|
|
;
|
|
; 12 registers per movem burst (d0-d7/a2-a5 = 48 bytes) is the maximum
|
|
; available: a0=src, a1=dst, a6=end sentinel. The row counter lives in the
|
|
; a1-vs-a6 compare rather than a d-register for exactly this reason.
|
|
; 512 = 10*48 + 32, hence ten 12-register bursts and one 8-register tail.
|
|
; Destination uses (d16,a1) displacement rather than post-increment because
|
|
; movem cannot post-increment a destination; the displacement costs 4 cycles
|
|
; per burst but saves an 8-cycle lea, so it is the cheaper of the two.
|
|
|
|
FLAG = $18000 ; 0 idle / 1 running / $FF done
|
|
VAR = $18004 ; variant selector, written by Lua
|
|
ITER = $18008 ; iteration count, written by Lua
|
|
SRCW = $60000 ; word-expanded frame 192*512 = 96KB
|
|
SRCB = $80000 ; byte-per-pixel frame 192*256 = 48KB
|
|
DST0 = $C08000 ; GVRAM + 32*1024 (first picture row)
|
|
DSTE = $C38000 ; GVRAM + 224*1024 (one past last)
|
|
|
|
org $10000
|
|
start:
|
|
move.l VAR.l,d0
|
|
move.l #1,FLAG.l ; timer starts here
|
|
cmp.l #1,d0
|
|
beq v1
|
|
cmp.l #2,d0
|
|
beq v2
|
|
cmp.l #4,d0
|
|
beq v4
|
|
bra v3
|
|
|
|
; ---------------------------------------------------------------- V1
|
|
v1: lea SRCW,a0
|
|
lea DST0,a1
|
|
lea DSTE,a6
|
|
v1row: movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,48(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,96(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,144(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,192(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,240(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,288(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,336(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,384(a1)
|
|
movem.l (a0)+,d0-d7/a2-a5
|
|
movem.l d0-d7/a2-a5,432(a1)
|
|
movem.l (a0)+,d0-d7
|
|
movem.l d0-d7,480(a1)
|
|
lea 1024(a1),a1
|
|
cmpa.l a6,a1
|
|
bne v1row
|
|
subq.l #1,ITER.l
|
|
bne v1
|
|
bra done
|
|
|
|
; ---------------------------------------------------------------- V2
|
|
v2: lea SRCB,a0
|
|
lea DST0,a1
|
|
lea DSTE,a6
|
|
v2row: move.w #255,d1
|
|
v2px: move.b (a0)+,d0
|
|
move.w d0,(a1)+ ; high byte is discarded by gvram_w
|
|
dbra d1,v2px
|
|
lea 512(a1),a1 ; skip the unused half of the line
|
|
cmpa.l a6,a1
|
|
bne v2row
|
|
subq.l #1,ITER.l
|
|
bne v2
|
|
bra done
|
|
|
|
; ---------------------------------------------------------------- V3
|
|
v3: lea SRCW,a0
|
|
movem.l (a0),d0-d7/a2-a5 ; load the burst once, outside the loop
|
|
lea DST0,a1
|
|
lea DSTE,a6
|
|
v3row: movem.l d0-d7/a2-a5,(a1)
|
|
movem.l d0-d7/a2-a5,48(a1)
|
|
movem.l d0-d7/a2-a5,96(a1)
|
|
movem.l d0-d7/a2-a5,144(a1)
|
|
movem.l d0-d7/a2-a5,192(a1)
|
|
movem.l d0-d7/a2-a5,240(a1)
|
|
movem.l d0-d7/a2-a5,288(a1)
|
|
movem.l d0-d7/a2-a5,336(a1)
|
|
movem.l d0-d7/a2-a5,384(a1)
|
|
movem.l d0-d7/a2-a5,432(a1)
|
|
movem.l d0-d7,480(a1)
|
|
lea 1024(a1),a1
|
|
cmpa.l a6,a1
|
|
bne v3row
|
|
subq.l #1,ITER.l
|
|
bne v3
|
|
bra done
|
|
|
|
; ---------------------------------------------------------------- V4
|
|
v4: lea SRCW,a0
|
|
lea DST0,a3 ; base of the current block row
|
|
lea DSTE,a4 ; one past the last block row
|
|
v4brow: move.l a3,a1
|
|
lea 512(a3),a5 ; 64 blocks * 8 bytes
|
|
v4blk: movem.l (a0)+,d0-d7 ; 32 bytes = one 4x4 block, expanded
|
|
movem.l d0-d1,(a1)
|
|
movem.l d2-d3,1024(a1)
|
|
movem.l d4-d5,2048(a1)
|
|
movem.l d6-d7,3072(a1)
|
|
addq.l #8,a1
|
|
cmpa.l a5,a1
|
|
bne.s v4blk
|
|
lea 4096(a3),a3 ; next block row is 4 picture lines
|
|
cmpa.l a4,a3
|
|
bne v4brow
|
|
subq.l #1,ITER.l
|
|
bne v4
|
|
bra done
|
|
|
|
done: move.l #$FF,FLAG.l ; timer stops here
|
|
halt: bra.s halt
|