Files
Dragon-s-Lair-X68k/src/player/load.i
T
prosolis 2676f3b835 Put the ring on the 68000, and find the disc stops whenever the player is not asking
ROADMAP P5. The loader moved in session 21 and the frame clock in 22; the ring
producer was the last policy living outside the machine. src/player/ring.i does
`aligned` placement, the descriptor ring, a prefill, 51.2's slack rule and a
seek, and the host keeps only the transport.

It needed a container change. `aligned` asks whether the next record fits
before the end of the ring -- a length asked BEFORE the record is fetched -- and
every reader in this tree answered that by walking the frame stream, which is
exactly what a player streaming off a disc cannot do. DLX4 carries nframes u16
record lengths in the scene header. Frame payloads are byte-identical to the
DLX3 encode, so no fitted constant moves; the scene header goes 5,920 to 6,164 B.

The producer reproduces the host's tiling exactly: 18 wraps, 14.7 KB mean hole,
pixel-exact, a third independent implementation of the same policy.

What it exposed is bigger than the item. A channel only moves bytes while it has
a request and only the CPU can issue one, so the disc stands still between
records by an amount the PLAYER sets, not the medium -- and no host-filled run
could see it. At 488 KB/s in a 256 KB ring a one-deep request queue gives away
6.8% of the pipe and underruns 59 of 120 frames; two-deep gives away 3.4% and
underruns none. The container's whole surplus over the wire is 8.7%, so the
player's own loop was spending most of the slack a branch point saves up.
Prefill is the weaker lever: six records of it still leaves 24 underruns.

Three silent bugs are recorded in FINDINGS 55.7 -- all produced wrong pixels or
a desync rather than a fault -- plus a rig one: MAME renders a screen line by
line, so snapshotting the frame the decoder finished in captures a tear that
reads exactly like a decoder bug.

check.sh gains the machine-owned ring and a seek with the decode after it.
decode.bin is unchanged at 1,296 B and a host-filled run executes none of the
new code, so every FINDINGS 49/51 figure stands. ALL GREEN before and after.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
2026-08-24 21:45:05 -07:00

257 lines
10 KiB
OpenEdge ABL

; ---------------------------------------------------------------------------
; load.i -- the two LOAD-TIME transforms, on the 68000 itself. ROADMAP P1+P2.
;
; Until now both of these were done host-side, in tools/bench/dlxload.py, and
; the rigs pushed the RESULT into emulated RAM. That was the right call while
; the inner loop was the thing being measured -- charging a once-per-scene cost
; to the per-frame path would have flattered or damned it for no reason -- but
; a player has no host. These are the bytes that replace it.
;
; The reference is tools/bench/dlxload.py and it stays the reference: this code
; is gated BYTE-FOR-BYTE against it (tools/bench/verify_load.py), palette words
; and darkest-entry index included. If the two ever disagree, the symptom in a
; rig would be wrong colours rather than a crash, which is exactly the class of
; bug the split was made to prevent.
;
; WHAT IT READS. The RAW container as it comes off the disc. The DLX header is
; fixed-layout and big-endian (tools/encoder/dlx.py):
; +0 magic 'DLX3' +12 k1 u16 +16 off_pal u32
; +4 W u16 +14 k4 u16 +20 off_cb1 u32
; +6 H u16 +24 off_cb4 u32
; +8 fps u16 +28 off_frm u32
; +10 nframes u16
; The three offsets are container-relative, so every one of them is an add of
; the base the loader was handed. Nothing here parses a frame record.
;
; WHAT IT WRITES. CB1 (8 KB) and CB4 (2 KB) expanded to one WORD per pixel at
; the addresses geom.i names, and 256 packed palette words straight into the
; graphics palette at $E82000. It also reports the darkest entry, which is what
; the letterbox is filled with until the encoder reserves a black one (23.4,
; still open).
;
; WHY WORD-PER-PIXEL. The block loop movems codebook entries straight into
; GVRAM with no unpacking, and the high byte of a GVRAM word write is discarded
; by the hardware, so the high byte is left zero and never has to be cleared.
; It also makes index scaling a shift rather than a multiply (lsl.w #5 / #3).
;
; SCRATCH. Three tables, built here and dead the moment the palette is packed:
; P6TAB 64 B 6-bit level -> the 8-bit value the hardware renders it as
; SQTAB 256 B the square of that, so the darkest-entry search has no muls
; DTAB 512 B err(v, I=0) - err(v, I=1) per 8-bit channel value, signed
; DTAB is what turns P2's per-entry minimum-squared-error choice of the shared
; LSB into three table reads and a sign test. Choosing I per entry rather than
; fixing it is worth 1.96 dB (FINDINGS 23.3), and it is a per-ENTRY decision
; across three channels, so it cannot be folded into a per-channel table alone.
; ---------------------------------------------------------------------------
LFLAG = $18040 ; 0 idle / 1 running / $FF done / $EE bad header
LHDR = $18044 ; -> raw container base
LDARK = $18048 ; <- index of the darkest palette entry
LK1 = $1804C ; <- k1, as the 68000 read it out of the header
LK4 = $18050 ; <- k4
LMODE = $18054 ; bit0 codebooks, bit1 palette entries,
; bit2 the three scratch tables
LITER = $18058 ; repeat count, so a 55 Hz host clock can time it
P6TAB = $19000 ; 64 bytes
SQTAB = $19040 ; 64 longs
DTAB = $19140 ; 256 words
GPAL = $E82000 ; graphics palette, 256 words
; ---------------------------------------------------------------- do_load
; in: a0 = container base, d1 = mode bits: 1 codebooks, 2 palette entries,
; 4 the scratch tables. A player builds the tables ONCE at boot (they
; describe the hardware's colour rendering and nothing about the scene) and
; then loads each scene with 3.
; out: d0 = 0 ok, -1 not a DLX3/DLX4 container. a0-a4 clobbered, a5 = base.
;
; The magic is accepted as 'DLX' plus a version byte of '3' OR ABOVE rather than
; as one constant. DLX4 (ROADMAP P5) adds the per-record index and a fifth
; header offset at +32; every field this routine reads is at its DLX3 place, so
; the transforms are version-independent and the check should be too. A version
; this loader has never seen is still refused -- '3' or above, not "anything
; that begins DLX".
do_load:
movea.l a0,a5
move.l (a5),d0
andi.l #$FFFFFF00,d0
cmpi.l #$444C5800,d0 ; 'DLX'
bne .bad
cmpi.b #'3',3(a5) ; ... version 3 or above
bcs .bad
move.w 12(a5),d0
ext.l d0
move.l d0,LK1.l
move.w 14(a5),d0
ext.l d0
move.l d0,LK4.l
btst #2,d1
beq.s .notab
move.l d1,-(sp)
bsr pal_tables
move.l (sp)+,d1
.notab:
btst #0,d1
beq.s .nocb
moveq #0,d2 ; the count is built as a LONG and the
move.w 12(a5),d2 ; high word must not carry junk into it
lsl.l #4,d2 ; k1 entries x 16 source bytes
movea.l 20(a5),a0
adda.l a5,a0
lea CB1,a1
bsr expand
moveq #0,d2
move.w 14(a5),d2
lsl.l #2,d2 ; k4 entries x 4 source bytes
movea.l 24(a5),a0
adda.l a5,a0
lea CB4,a1
bsr expand
.nocb:
btst #1,d1
beq.s .nopal
bsr pal_pack
.nopal:
moveq #0,d0
rts
.bad: moveq #-1,d0
rts
; ---------------------------------------------------------------- expand
; One source byte -> one destination word, high byte zero.
; in: a0 src, a1 dst, d2 = source byte count. Always a multiple of 4: CB1 is
; k1*16 and CB4 is k4*4, so no remainder case can exist and none is written.
; A junk high word here is not a slow path, it is a WRONG one: `lsr.l #2` walks
; two of its bits down into the low word and the dbra count comes out long.
expand:
lsr.l #2,d2
subq.l #1,d2 ; k<=256, so the count fits a dbra
moveq #0,d0
.e1: move.b (a0)+,d0
move.w d0,(a1)+
move.b (a0)+,d0
move.w d0,(a1)+
move.b (a0)+,d0
move.w d0,(a1)+
move.b (a0)+,d0
move.w d0,(a1)+
dbra d2,.e1
rts
; ---------------------------------------------------------------- pal_tables
; The three scratch tables. SCENE-INDEPENDENT, every one of them: they describe
; how the CRTC renders a 5-bit channel plus the shared LSB, which is a property
; of the machine. A player builds them once at boot and never again, which is
; why they are a separate entry point rather than the head of pal_pack -- see
; FINDINGS 53.3 for what that is worth.
pal_tables:
; -- P6TAB[x] = ((x<<2)|(x>>4)) & $FF, and SQTAB[x] = P6TAB[x]^2
lea P6TAB,a0
lea SQTAB,a1
moveq #0,d1
.p1: move.w d1,d0
lsl.w #2,d0
move.w d1,d2
lsr.w #4,d2
or.w d2,d0
andi.w #$FF,d0
move.b d0,(a0)+
move.w d0,d2
mulu d2,d2
move.l d2,(a1)+
addq.w #1,d1
cmpi.w #64,d1
bne.s .p1
; -- DTAB[v] = (render(v,0)-v)^2 - (render(v,1)-v)^2, signed
lea P6TAB,a0
lea DTAB,a1
moveq #0,d1
.p2: move.w d1,d2
lsr.w #2,d2
andi.w #$3E,d2 ; x0 = (v>>3)<<1
moveq #0,d3
move.b 0(a0,d2.w),d3
sub.w d1,d3
muls d3,d3
moveq #0,d4
move.b 1(a0,d2.w),d4
sub.w d1,d4
muls d4,d4
sub.l d4,d3
move.w d3,(a1)+
addq.w #1,d1
cmpi.w #256,d1
bne.s .p2
rts
; ---------------------------------------------------------------- pal_pack
; 24-bit RGB -> GGGGGRRRRRBBBBBI, the shared LSB chosen per entry by minimum
; squared error, written to the palette registers. Identical arithmetic to
; dlxload.pack_palette, including its tie-breaks: I stays 0 when the two errors
; are equal, and the darkest entry is the FIRST index at the minimum.
; in: a5 = container base, and pal_tables already run.
pal_pack:
movea.l 16(a5),a0
adda.l a5,a0 ; -> 256 x RGB888
lea GPAL,a1
lea DTAB,a2
lea SQTAB,a4 ; P6TAB is not needed here: the rendered
; value is only ever wanted SQUARED
move.l #$7FFFFFFF,d6
clr.l LDARK.l
moveq #0,d7
.p3: moveq #0,d1
move.b (a0)+,d1 ; R
moveq #0,d2
move.b (a0)+,d2 ; G
moveq #0,d3
move.b (a0)+,d3 ; B
move.w d1,d0
add.w d0,d0
move.w 0(a2,d0.w),d4
move.w d2,d0
add.w d0,d0
add.w 0(a2,d0.w),d4
move.w d3,d0
add.w d0,d0
add.w 0(a2,d0.w),d4 ; sum of err0-err1 over the three
moveq #0,d5
tst.w d4
ble.s .p4
moveq #1,d5 ; I=1 only when it is STRICTLY better
.p4: lsr.w #3,d1 ; fR
lsr.w #3,d2 ; fG
lsr.w #3,d3 ; fB
move.w d2,d4
lsl.w #5,d4
or.w d1,d4
lsl.w #6,d4 ; (fG<<11)|(fR<<6)
move.w d3,d0
add.w d0,d0
or.w d0,d4
or.w d5,d4
move.w d4,(a1)+ ; -> the palette register
add.w d1,d1 ; x = (f<<1)|I, per channel
or.w d5,d1
add.w d2,d2
or.w d5,d2
add.w d3,d3
or.w d5,d3
lsl.w #2,d1 ; SQTAB holds longs
move.l 0(a4,d1.w),d0
lsl.w #2,d2
add.l 0(a4,d2.w),d0
lsl.w #2,d3
add.l 0(a4,d3.w),d0 ; squared distance from black
cmp.l d6,d0
bge.s .p5
move.l d0,d6
move.l d7,LDARK.l ; first index at the minimum wins
.p5: addq.w #1,d7
cmpi.w #256,d7
bne .p3
rts