; DLX1 frame decoder for a stock 68000 @ 10MHz -- direct-to-GVRAM, single path. ; ; WHY ONE PATH. FINDINGS 24.5/25.6 specified two display paths chosen per ; frame (compose-in-RAM-then-blit vs decode-direct-to-GVRAM, crossing at 70% of ; blocks changed). That plan is incoherent: the compose path assembles a FULL ; frame in RAM so its blit can be row-linear, and the pixels it does not decode ; this frame -- the SKIP blocks -- can only come from a RAM copy of the previous ; reconstruction, which the direct path deliberately never writes. Every direct ; frame invalidates the next compose frame's reference. Measured on the Singe ; window: 70 of 120 frames display stale pixels, worst frame 18.8% of the ; screen. See tools/analysis/10_pathmix_drift.py and FINDINGS 28. ; ; Every coherent version of the mix is worse than plain direct on the median, ; so this decoder implements direct only. It also drops the 96KB RAM reference ; frame entirely: the previous frame is already in GVRAM, so SKIP is genuinely ; free -- no read, no write, just a pointer advance. ; ; GEOMETRY (tools/bench/crtc_mode.lua). 256-colour page: one pixel per WORD of ; CPU address space, 1024-byte line stride, picture in rows 32..223. A 4x4 ; block is therefore 4 rows of 8 bytes at a 1024-byte stride, so displacements ; 0/1024/2048/3072 all fit a 16-bit offset and one base pointer covers a block. ; A block row is 4 picture rows = 4096 bytes; 64 blocks x 8 bytes = 512. ; ; The HIGH byte of every GVRAM word write is discarded by the hardware (MAME ; 0.277 x68k_crtc.cpp:501, gvram_w case 0x0100), which is what makes the RAW ; path cheap: it never has to clear the odd bytes it builds. ; ; CODEBOOKS are pre-expanded to word-per-pixel form by the loader, so the inner ; loop movems them straight out with no unpacking: ; CB1 entry = 16 words, row-major = 32 bytes (k1=256 -> 8KB) ; CB4 entry = 4 words, row-major (2x2) = 8 bytes (k4=256 -> 2KB) ; Index scaling is therefore a shift, not a multiply: lsl.w #5 and lsl.w #3. ; ; REGISTERS are fully committed -- a0 payload, a1 mode header, a2/a3 codebooks, ; a4 block cursor, a5 block-row end, a6 block-row base, d0-d7 the 32-byte block ; transfer. Nothing survives a block, which is why each block re-reads its mode ; from (a1) rather than holding the packed byte in a register. That re-read is ; paid only by blocks that are NOT all-SKIP: a header byte of zero clears four ; blocks with one tst.b, and SKIP is the median block. ; ; LITERAL SPANS (v7, FINDINGS 40). A run of horizontally adjacent dirty blocks ; is cheaper to paint as four ROW-LINEAR runs of word-expanded literal pixels ; than as blocks: 226 clocks per 4x4 block at a run of 4, against V1's 299.9, ; and the break-even is a run of 2. The run's blocks read SKIP in the mode ; header and the span section paints them instead, so the block loop below is ; unchanged -- it sees a SKIP and advances, exactly as it does for a genuinely ; held block. ; ; The section sits BETWEEN the mode header and the block payload because that is ; the only place the 68000 can reach without first parsing something of variable ; length: the header is a fixed 768 bytes. Per span the record is {u32 absolute ; GVRAM address, u16 coarse displacement}, then the coarse pixels, then a u16 ; FINE displacement, then the fine pixels. ; ; The two displacements are jumps into two unrolled copy chains -- 24 pixels per ; coarse unit (a 12-register movem pair) and 2 per fine unit (one ; `move.l (a0)+,(a2)+`) -- so a span of any length is straight-line code with no ; loop, no remainder and no address arithmetic. A run of 4x4 blocks is always a ; multiple of 4 pixels long, and 4 is a multiple of the 2-pixel fine quantum, so ; NOTHING is padded (FINDINGS 40.3). ; ; The fine displacement is in the STREAM rather than in the span record because ; that is what pays for the second dispatch: when the coarse chain falls out ; into `move.w (a0)+,d0 / jmp`, d0 is dead payload and a0 is already pointing at ; it, so the decoder holds nothing extra across the copy and keeps all twelve ; payload registers (FINDINGS 40.4). Twelve is why the coarse unit is 24 pixels ; and not V5's 16, and it is the whole reason the per-pixel cost is 9.143 rather ; than 10.459 (FINDINGS 30.4). ; ; a1 (the mode header cursor) is one of those twelve, so it goes on the stack ; across the span pass. Two long accesses per frame, against the 24 pixels a ; register buys per chain unit. ; ; ALIGNMENT. Frame records are [u32 length][768-byte mode header][payload] laid ; end to end, and payload lengths are arbitrary -- so record boundaries land on ; odd addresses, and `move.l (a0)+,d0` on an odd address is an ADDRESS ERROR on ; a 68000, not a slow read. It killed the first run: frame 0 decoded perfectly, ; then the length read for frame 1 at $03220F vectored into the IPL. Being ; big-endian is only half of what "the 68000 reads it with a plain move" needs. ; This decoder rounds each record start up to 4; the container itself should ; carry the padding so a streaming player can DMA records straight into place. ; FINDINGS 28.3. FLAG = $18000 ; 0 idle / 1 running / $FF done / $EE desync ITER = $18008 ; outer repeat count, written by Lua NFR = $1800C ; frames per pass FPTR = $18010 ; -> first frame record SCR_N = $18014 ; frames remaining this pass SCR_END = $18018 ; expected end of the current payload include "src/player/geom.i" org $10000 start: move.l #1,FLAG.l ; timer starts here outer: move.l FPTR.l,a0 move.l NFR.l,SCR_N.l frameloop: move.l (a0)+,d0 ; u32 payload length, big-endian lea 0(a0,d0.l),a1 move.l a1,SCR_END.l ; where the payload must end move.l a0,a1 ; a1 = packed mode header lea MODEB(a0),a0 ; a0 = span section bsr paint_spans ; -> a0 = block payload, a1 preserved bsr decode_frame cmpa.l SCR_END.l,a0 ; bitstream desync is silent otherwise bne desync move.l a0,d0 ; next record starts on a 4-byte boundary addq.l #3,d0 and.b #$FC,d0 move.l d0,a0 subq.l #1,SCR_N.l bne frameloop subq.l #1,ITER.l bne outer move.l #$FF,FLAG.l ; timer stops here hold: bra.s hold desync: move.l #$EE,FLAG.l bra.s hold include "src/player/frame.i"