Put the player on a real volume, and find the write window is the frame
ROADMAP K3. src/player/packed.s (2,898 B) brings up its own display, builds
its own 193-entry DMA chain, keeps its own frame clock off V-DISP and fetches
every record itself with READ(10) off a CZ-6BS1. The rig writes no picture
byte, no palette entry and no CRTC register.
120 of 120 frames pixel-exact, every one compared, in both palette orders --
the gate had to grow to do it, because a packed frame is a LITERAL and the
codec's recursion was what made one comparison audit 120.
And the write window turns out to be the frame. A packed write requires R20
bit 11, buffer mode blanks the layer, and a DMAC-direct player holds the
window open for the whole data phase, so
dark fraction of a slot = record bytes / (DATA-PHASE rate x slot)
which is 1.0 at the container's own 582.0 KB/s: every frame delivered, on
time, pixel-exact, and none of them displayed. The rate in that expression is
the BURST rate, a third hardware number B1 has no test for. It reverses 61.5's
ranking -- a packed player that DMAs to RAM and paints with the measured 27.3%
blit is on screen 72.7% of every slot at any rate, and the two are equal only
at 2,131 KB/s = 3.7x the wire.
And a held channel costs the frame clock half its ticks without the clock
being able to tell: 487 of 1,038 V-DISP edges lost, zero late frames reported,
the player believing 12 fps while the screen ran at 6.37.
FINDINGS 64. ROADMAP K4 opened and fenced behind B2.
check.sh ALL GREEN before and after.
Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -0,0 +1,512 @@
|
||||
; ---------------------------------------------------------------------------
|
||||
; packed.s -- THE PLAYER WITH NO DECODER, END TO END, OFF THE DISC. ROADMAP K3.
|
||||
;
|
||||
; WHAT THE OTHER FRONT-ENDS IN THIS DIRECTORY DO, AND WHY THIS ONE IS DIFFERENT.
|
||||
; decode.s parses a DLX record and draws 4x4 blocks; stream.s does the same out
|
||||
; of a bounded ring that ring.i fills through xfer.i. Both are CODECS: a record
|
||||
; is a program the 68000 executes against GVRAM, and every cost table in
|
||||
; docs/FINDINGS.md from 24 to 45 is a cost of executing it. FINDINGS 61 priced
|
||||
; the alternative and it won: at the 9 clk/B dual-address floor the shipping
|
||||
; codec is 110.4% of a 12 fps frame and a decoder-free packed literal frame is
|
||||
; 55.2%. Decoding 37,585 bytes costs more than not decoding 49,152.
|
||||
;
|
||||
; So THERE IS NO DECODER BELOW. There is no ring either, and that is the same
|
||||
; fact rather than a second one: a ring exists because DLX records are
|
||||
; variable-length and the block loop needs the next one whole and contiguous
|
||||
; (49.2). A packed record's length is GEOMETRY -- 512 B of palette and 192 rows
|
||||
; of 256 B, 97 sectors exactly -- so record `i` is at LBA0 + i*97 and there is
|
||||
; nothing to place, nothing to wrap and nothing to index (tools/encoder/dlxp.py).
|
||||
;
|
||||
; WHAT THE 68000 DOES PER FRAME, IN FULL:
|
||||
;
|
||||
; 1. wait for its own frame tick (src/player/clock.i, off the CRTC's V-DISP)
|
||||
; 2. set CRTC R20 bit 11 -- the GVRAM write window
|
||||
; 3. issue one READ(10) whose DATA IN phase is an HD63450 channel walking a
|
||||
; 193-entry chain: the palette registers at $E82000, then 192 GVRAM rows a
|
||||
; 1,024 B line stride apart (FINDINGS 62)
|
||||
; 4. clear R20 bit 11
|
||||
;
|
||||
; That is the whole video path. There is no per-frame PAINT: between the
|
||||
; channel start and the channel's COC the 68000 executes nothing at all in the
|
||||
; held configuration, and in the stealing one it executes only its own wait
|
||||
; loop. The array is SCENE-CONSTANT -- the packed layout spends both 256-colour
|
||||
; pages, so there is no page to flip and no destination that changes.
|
||||
;
|
||||
; THE TWO THINGS THIS FILE EXISTS TO FIND OUT, neither of which any earlier run
|
||||
; could ask:
|
||||
;
|
||||
; a. DOES A CHAINED TRANSFER RUN BACK TO BACK AT 12 fps? 61.7.2 named this as
|
||||
; the specific untested thing behind K3's large simplification, and it is
|
||||
; the reason the frame clock is in here rather than a host tick: a chain
|
||||
; that has to be restarted 120 times in ten seconds is a different claim
|
||||
; from one that ran once in a gate.
|
||||
; b. WHAT DOES HOLDING THE BUS COST THE CLOCK? A held channel halts the
|
||||
; 68000, and the frame clock is an INTERRUPT off V-DISP. Edges that fall
|
||||
; while the CPU is halted are not counted twice by the MFP -- the pending
|
||||
; bit is one bit -- so a transfer long enough to span two V-DISPs makes the
|
||||
; player's own clock LOSE TIME. Nothing in this project has run a transfer
|
||||
; and a clock at once before, so nothing could have seen it. CLK_VDISP is
|
||||
; the machine's count and the host's raster count is the truth; the gate
|
||||
; compares them rather than trusting either.
|
||||
;
|
||||
; BOTH CONFIGURATIONS ARE SELECTABLE (PG_HELD) FOR EXACTLY THAT REASON. 59.3
|
||||
; already showed that an auto-requested channel is charged by TIME rather than
|
||||
; by byte, so "held" and "stealing" are not two speeds of the same thing: held
|
||||
; is the CPU stopped for as long as the record takes to arrive, and stealing is
|
||||
; the CPU running against a channel taking its share. A player has to keep a
|
||||
; clock, read a joystick and feed ADPCM, so which of the two is survivable is a
|
||||
; design question and not a benchmark.
|
||||
;
|
||||
; AND IT IS STILL NOT A RATE. MAME's HD63450 is configured in wall-clock
|
||||
; attotimes (42.5) and its held mode halts the CPU rather than charging it
|
||||
; cycles per operand, so nothing below measures `W`. What it measures is
|
||||
; whether the SHAPE works: one start, 193 destinations, 120 times, on a clock
|
||||
; the machine keeps itself, with every frame pixel-exact off a real volume.
|
||||
; ---------------------------------------------------------------------------
|
||||
|
||||
; ---- inputs, written by the rig before the CPU is launched
|
||||
PG_FLAG = $18900 ; 0 idle / 1 running / $FF done / $Exx failed
|
||||
PG_NFR = $18904 ; frames in the scene
|
||||
PG_FPS = $18908 ; frame rate the clock is asked for
|
||||
PG_LBA0 = $1890C ; LBA of record 0. A WORD AND NOT A CONSTANT:
|
||||
; a shipping volume has a filesystem in front of
|
||||
; the stream, and this is the one number that
|
||||
; changes when it does (xfer.i says the same).
|
||||
PG_RECS = $18910 ; sectors in a record -- 97, and the container
|
||||
; is what says so
|
||||
PG_PALL = $18914 ; 1 = the palette is LAST in the record. It is
|
||||
; a CONTAINER property (dlxp.py flags bit 1) and
|
||||
; the player reads it rather than assuming it:
|
||||
; FINDINGS 62.5/63.4 priced both orders at
|
||||
; -12.8 dB for one paint and could not choose
|
||||
; between them, so K3 runs both.
|
||||
PG_HELD = $18918 ; 1 = DM_HELD_* (burst, bus held), 0 = DM_STEAL_*
|
||||
PG_PACEON = $1891C ; 1 = obey the frame clock. 0 free-runs, which
|
||||
; tests the CHAIN without the clock in the way.
|
||||
PG_ITER = $18920 ; passes over the scene; >1 exercises the SEEK,
|
||||
; which for this container is arithmetic
|
||||
|
||||
; ---- outputs
|
||||
PG_SHOWN = $18930 ; frames displayed. Bumped AFTER bit 11 is
|
||||
; cleared, so a rig that snapshots on a change
|
||||
; is snapshotting a frame that is on screen.
|
||||
PG_ERR = $18934 ; SC_ERR of the first failed read, 0 = none
|
||||
PG_ERRAT = $18938 ; ...and the frame it failed on
|
||||
PG_LATE = $1893C ; frames that reached the gate with their tick
|
||||
; already past -- the previous frame overran
|
||||
PG_LATE1 = $18940 ; the first of them
|
||||
PG_LATEM = $18944 ; the worst, in whole ticks
|
||||
PG_VDISP = $18948 ; CLK_VDISP as the machine counted it
|
||||
PG_VD0 = $1894C ; ...and as it stood when frame 0 started, so a
|
||||
; rig can charge only the frames it ran
|
||||
PG_TSPIN = $18950 ; total DM_SPIN over the run: the CPU's own trips
|
||||
; round the transfer wait. Held, this is one per
|
||||
; frame and the CPU did nothing else all scene.
|
||||
PG_GSPIN = $18954 ; total pace-gate polls: what the CPU had LEFT
|
||||
PG_LOSTV = $18958 ; V-DISP edges the machine did not see, summed
|
||||
; per frame -- see pg_frame
|
||||
PG_ARRN = $1895C ; entries the array was built with (an assertion
|
||||
; the rig reads back rather than a comment)
|
||||
|
||||
; ---- the array chain. 193 x 6 B = 1,158 B, built once at scene setup.
|
||||
; $1B000 and not $19000: dmagate.s's two arrays live at $19000/$19100 and
|
||||
; ring.i's disc-offset table at $19400, and a front-end that shares an address
|
||||
; with another front-end is how DM_USE landed on ring.i's mailbox (dma.i).
|
||||
PG_ARR = $1B000
|
||||
PG_MAXE = 256 ; the chain cannot be longer than this
|
||||
|
||||
; ---- geometry. The container's, not the screen's: the screen is 256x256 and
|
||||
; the picture is 256x192, so the difference is letterbox and is STATIC SETUP.
|
||||
PG_W = 256
|
||||
PG_H = 192
|
||||
PG_ROWB = PG_W/2 ; 128 words = 256 BYTES a row. THE WHOLE POINT:
|
||||
; 1.0 B/pixel, because bit 11 stops GVRAM
|
||||
; masking the high byte away and the two
|
||||
; 256-colour pages are scrolled 384 apart
|
||||
; (FINDINGS 46.5/47.1, dlxp.py).
|
||||
PG_STRIDE = 1024 ; GVRAM line stride, in bytes
|
||||
PG_PALB = 512 ; 256 GRB555+I words -- and exactly one sector
|
||||
|
||||
SCRW = 256
|
||||
SCRH = 256
|
||||
GVRAM = $C00000
|
||||
GPAL = $E82000
|
||||
VC0 = $E82400
|
||||
VC1 = $E82500
|
||||
VC2 = $E82600
|
||||
CONTRAST = $E8E001
|
||||
; The graphic scroll registers, named rather than written as CRTC+n*2.l: a
|
||||
; 256-colour page is assembled from TWO nibble planes with independent scroll
|
||||
; registers (px68k Grp_DrawLine8 reads scroll sets page*2 and page*2+1), so both
|
||||
; of a page's registers have to agree or the page tears between its low and high
|
||||
; nibble. Naming them is what makes the pairing visible at the write.
|
||||
CR_P0X0 = CRTC+12*2
|
||||
CR_P0Y0 = CRTC+13*2
|
||||
CR_P0X1 = CRTC+14*2
|
||||
CR_P0Y1 = CRTC+15*2
|
||||
CR_P1X0 = CRTC+16*2
|
||||
CR_P1Y0 = CRTC+17*2
|
||||
CR_P1X1 = CRTC+18*2
|
||||
CR_P1Y1 = CRTC+19*2
|
||||
PG_YOFF = (SCRH-PG_H)/2 ; 32
|
||||
PG_TOP = GVRAM+PG_YOFF*PG_STRIDE
|
||||
PG_BLACK = $FFFF ; letterbox: index 255 in BOTH bytes. NOT 0 --
|
||||
; index 0 is page 1's transparency key and black
|
||||
; lives at 255 (dlxp.py, vq.frame_palette).
|
||||
|
||||
; R20. bit 11 = G-VRAM set to buffer, i.e. the WRITE WINDOW; bits 9-8 = 01
|
||||
; 256 colours; bit 4 = 31.5 kHz; bits 3-0 = 256 lines, 256 dots. The two
|
||||
; values differ in EXACTLY bit 11, which is what makes the window a window.
|
||||
PG_R20D = $0110 ; displaying
|
||||
PG_R20B = $0910 ; the write window open
|
||||
|
||||
include "src/player/geom.i"
|
||||
|
||||
org $10000
|
||||
start:
|
||||
move.l #1,PG_FLAG.l
|
||||
clr.l PG_SHOWN.l
|
||||
clr.l PG_ERR.l
|
||||
move.l #-1,PG_ERRAT.l
|
||||
clr.l PG_LATE.l
|
||||
move.l #-1,PG_LATE1.l
|
||||
clr.l PG_LATEM.l
|
||||
clr.l PG_TSPIN.l
|
||||
clr.l PG_GSPIN.l
|
||||
clr.l PG_LOSTV.l
|
||||
|
||||
; ---- 1. the display. THE PLAYER OWNS IT, and that is a change from every
|
||||
; other rig in this tree: tools/bench/stream.lua and decode.lua call
|
||||
; crtc_mode.lua's MODE.apply from the host, which is fine for a decoder gate and
|
||||
; is not a player. A player boots into a machine the IPL left in 768x512
|
||||
; 16-colour and has to get to 256x256 256-colour packed by itself, and the
|
||||
; scroll registers and the priority word are load-bearing rather than cosmetic:
|
||||
; 47.5 measured VC1 = $0000 putting page 0 on top and blacking the right half of
|
||||
; the screen.
|
||||
bsr pg_video
|
||||
|
||||
; ---- 2. the static half of the picture (FINDINGS 47.2). Words 128..511 of
|
||||
; every row, and the letterbox rows entire, are written ONCE and never again --
|
||||
; which is what makes the per-frame payload exactly the picture. Page 1's
|
||||
; storage at 384..511 is what the +384 scroll puts under screen columns 0..127,
|
||||
; and it must read 0 so that the opaque page 0 shows through there.
|
||||
bsr pg_static
|
||||
|
||||
; ---- 3. the chain the channel will walk, 193 entries of it (FINDINGS 62).
|
||||
bsr pg_array
|
||||
|
||||
; ---- 4. the transport. ONCE, not per frame: scsi_init resets the SPC and
|
||||
; clears DM_USE, so a per-frame call would quietly put the data phase back on
|
||||
; the CPU and the run would still deliver every byte -- 87 clocks each (58).
|
||||
bsr scsi_init
|
||||
bsr pg_dmacfg
|
||||
|
||||
; ---- 5. the clock. AFTER the SPC, because clk_init lowers the interrupt mask
|
||||
; to $2500 and the bring-up should not be the first thing running with level 6
|
||||
; open. It also CLEARS PACE, so tick 0 is the instant the scene starts.
|
||||
move.l PG_FPS.l,CLK_FPS.l
|
||||
tst.l PG_PACEON.l
|
||||
beq.s .noclk
|
||||
bsr clk_init
|
||||
tst.l CLK_ERR.l
|
||||
beq.s .noclk
|
||||
move.l #$E1,PG_FLAG.l ; the mode is not one HFREQ describes
|
||||
bra pg_hold
|
||||
.noclk:
|
||||
move.l CLK_VDISP.l,PG_VD0.l
|
||||
|
||||
; ---- 6. the scene.
|
||||
move.l PG_ITER.l,d6
|
||||
tst.l d6
|
||||
bne.s .it
|
||||
moveq #1,d6
|
||||
.it:
|
||||
pg_pass:
|
||||
moveq #0,d7 ; d7 = frame index within the pass
|
||||
pg_loop:
|
||||
bsr pg_gate ; wait for tick d7
|
||||
bsr pg_frame ; and paint it
|
||||
tst.l d0
|
||||
bmi.s pg_failed
|
||||
addq.l #1,d7
|
||||
move.l PG_NFR.l,d0
|
||||
cmp.l d0,d7
|
||||
bcs.s pg_loop
|
||||
; ---- a pass boundary. For this container a SEEK IS ARITHMETIC: there is no
|
||||
; ring to discard, no index to walk and no prefill to climb, so the next pass's
|
||||
; first record is simply LBA0 again. That is the whole of what K3 deletes, and
|
||||
; it is worth one line of code and four of comment because it is the largest
|
||||
; simplification in the project (ROADMAP K3) and it does not look like one.
|
||||
subq.l #1,d6
|
||||
bne.s pg_pass
|
||||
|
||||
move.l CLK_VDISP.l,PG_VDISP.l
|
||||
bsr clk_stop
|
||||
move.l #$FF,PG_FLAG.l
|
||||
bra.s pg_hold
|
||||
pg_failed:
|
||||
move.l CLK_VDISP.l,PG_VDISP.l
|
||||
bsr clk_stop
|
||||
move.l #$E2,PG_FLAG.l
|
||||
pg_hold:
|
||||
bra.s pg_hold
|
||||
|
||||
; ---------------------------------------------------------------- pg_gate
|
||||
; Frame d7 may not START before tick d7 (the rule is stream.s's, unchanged, and
|
||||
; deliberately so: the same gate against the same clock is what makes a paced
|
||||
; packed run comparable with a paced codec run). A frame that arrives with its
|
||||
; tick ALREADY past did not idle for a single poll, which means the previous
|
||||
; frame used its whole slot and then some -- that is the underrun this player
|
||||
; can have, and it is counted rather than absorbed.
|
||||
pg_gate:
|
||||
movem.l d0-d1,-(sp)
|
||||
tst.l PG_PACEON.l
|
||||
beq.s .out
|
||||
move.l CLK_PACE.l,d0
|
||||
cmp.l d7,d0
|
||||
bcs.s .wait ; PACE < d7: early, the common case
|
||||
tst.l d7
|
||||
beq.s .out ; frame 0 starts AT tick 0 by definition
|
||||
tst.l PG_LATE.l
|
||||
bne.s .nf
|
||||
move.l d7,PG_LATE1.l
|
||||
.nf: addq.l #1,PG_LATE.l
|
||||
sub.l d7,d0 ; whole ticks overrun
|
||||
cmp.l PG_LATEM.l,d0
|
||||
bls.s .out
|
||||
move.l d0,PG_LATEM.l
|
||||
bra.s .out
|
||||
.wait: addq.l #1,PG_GSPIN.l
|
||||
move.l CLK_PACE.l,d0
|
||||
cmp.l d7,d0
|
||||
bcs.s .wait
|
||||
.out: movem.l (sp)+,d0-d1
|
||||
rts
|
||||
|
||||
; ---------------------------------------------------------------- pg_frame
|
||||
; ONE FRAME. Open the write window, hand the record to the channel, close it.
|
||||
; Returns d0 < 0 on a transport failure.
|
||||
;
|
||||
; THE WINDOW IS OPENED ROUND THE TRANSFER AND NOT ROUND THE SCENE, and that is
|
||||
; not tidiness. R20 bit 11 blanks the graphics layer while it is set (measured,
|
||||
; tools/bench/crtc_mode.lua) -- 47.4/ROADMAP B2 is exactly the question of
|
||||
; whether a real board does too -- so a scene-long window would show nothing at
|
||||
; all. Held round the transfer it is a shutter: the screen is dark for as long
|
||||
; as the record takes to land and shows a COMPLETE frame the rest of the time.
|
||||
; That is also why this player cannot tear the way FINDINGS 41's decoder does:
|
||||
; there is no instant at which a half-written picture is displayable.
|
||||
;
|
||||
; CLK_VDISP IS SAMPLED EITHER SIDE OF THE TRANSFER. The difference is the
|
||||
; V-DISP edges the machine SAW; the raster produced its own number regardless.
|
||||
; In the held configuration the CPU is stopped for the whole transfer, so any
|
||||
; edge that falls inside it is one the MFP can only remember once -- and a frame
|
||||
; clock built on counting them loses time it can never get back. The subtraction
|
||||
; here is what makes that visible without a host in the loop.
|
||||
pg_frame:
|
||||
movem.l d1-d5/a1,-(sp)
|
||||
move.l CLK_VDISP.l,d5
|
||||
|
||||
move.w #PG_R20B,CRTC_R20.l ; the write window opens
|
||||
|
||||
; LBA = PG_LBA0 + d7 * PG_RECS. Arithmetic, not a lookup: a packed record's
|
||||
; length is geometry, so this player carries no record index at all (dlxp.py).
|
||||
move.l d7,d3
|
||||
move.l PG_RECS.l,d0
|
||||
mulu d0,d3 ; frames * sectors, both small
|
||||
add.l PG_LBA0.l,d3
|
||||
move.l PG_RECS.l,d4
|
||||
lea GVRAM,a1 ; IGNORED under chaining -- the channel
|
||||
; takes MAR from the array's first entry
|
||||
; -- and passed so that this call site
|
||||
; reads like every other one in the tree
|
||||
bsr scsi_read
|
||||
move.l d0,d1
|
||||
|
||||
move.w #PG_R20D,CRTC_R20.l ; ...and closes. The frame is up.
|
||||
|
||||
move.l DM_SPIN.l,d0
|
||||
add.l d0,PG_TSPIN.l
|
||||
move.l CLK_VDISP.l,d0
|
||||
sub.l d5,d0 ; edges seen across the transfer
|
||||
beq.s .nolost
|
||||
subq.l #1,d0 ; one edge per frame slot is expected;
|
||||
add.l d0,PG_LOSTV.l ; what is counted is the SURPLUS, and a
|
||||
.nolost: ; deficit cannot be seen from in here --
|
||||
; the host's raster count is what says
|
||||
; how many there really were
|
||||
move.l d1,d0
|
||||
bmi.s .err
|
||||
addq.l #1,PG_SHOWN.l
|
||||
movem.l (sp)+,d1-d5/a1
|
||||
rts
|
||||
.err: tst.l PG_ERR.l
|
||||
bne.s .err2
|
||||
move.l SC_ERR.l,PG_ERR.l
|
||||
move.l d7,PG_ERRAT.l
|
||||
.err2: moveq #-1,d0
|
||||
movem.l (sp)+,d1-d5/a1
|
||||
rts
|
||||
|
||||
; ---------------------------------------------------------------- pg_dmacfg
|
||||
; The channel's configuration, and the choice between the only two rows of the
|
||||
; ladder MAME has a code path for (dma.i: no EXREQ wiring, no single-address
|
||||
; path, and only burst modelled as held). OCR gets CHAIN = %10 on top, which is
|
||||
; sequential array chaining and is what makes 193 destinations one start.
|
||||
pg_dmacfg:
|
||||
move.l #DM_STEAL_DCR,d0
|
||||
move.l #DM_STEAL_OCR,d1
|
||||
tst.l PG_HELD.l
|
||||
beq.s .set
|
||||
move.l #DM_HELD_DCR,d0
|
||||
move.l #DM_HELD_OCR,d1
|
||||
.set: move.l d0,DM_DCRV.l
|
||||
ori.l #$08,d1 ; OCR CHAIN = %10, array chain
|
||||
move.l d1,DM_OCRV.l
|
||||
move.l #PG_ARR,DM_BARV.l
|
||||
move.l PG_ARRN.l,DM_BTCV.l
|
||||
move.l #1,DM_USE.l ; AFTER scsi_init, which clears it
|
||||
rts
|
||||
|
||||
; ---------------------------------------------------------------- pg_array
|
||||
; The 193 {u32 MAR, u16 MTC} entries, built once. SCENE-CONSTANT: the packed
|
||||
; layout spends both 256-colour pages, so there is no page to flip and no
|
||||
; destination that changes from frame to frame (FINDINGS 62).
|
||||
;
|
||||
; THE ORDER IS THE CONTAINER'S. Palette first or 193rd is one paint of
|
||||
; mismatch either way and 63.4 priced both at -12.8 dB without being able to
|
||||
; choose; so the format records it (dlxp.py flags bit 1) and this reads it. An
|
||||
; array built the other way round from the record it is fed would not fail -- it
|
||||
; would paint 192 rows of picture into the palette registers and 512 B of
|
||||
; palette across the top two rows of the screen, which is a picture, and a gate
|
||||
; that only checked for errors would pass it.
|
||||
pg_array:
|
||||
movem.l d0-d2/a0-a1,-(sp)
|
||||
lea PG_ARR,a0
|
||||
tst.l PG_PALL.l
|
||||
bne.s .rows
|
||||
bsr .pal
|
||||
.rows:
|
||||
lea PG_TOP,a1
|
||||
move.w #PG_H-1,d0
|
||||
.r: move.l a1,(a0)+
|
||||
move.w #PG_ROWB*2,(a0)+
|
||||
lea PG_STRIDE(a1),a1
|
||||
dbra d0,.r
|
||||
tst.l PG_PALL.l
|
||||
beq.s .done
|
||||
bsr .pal
|
||||
.done:
|
||||
move.l a0,d0
|
||||
sub.l #PG_ARR,d0
|
||||
divu #6,d0
|
||||
andi.l #$FFFF,d0
|
||||
move.l d0,PG_ARRN.l
|
||||
movem.l (sp)+,d0-d2/a0-a1
|
||||
rts
|
||||
.pal: move.l #GPAL,(a0)+
|
||||
move.w #PG_PALB,(a0)+
|
||||
rts
|
||||
|
||||
; ---------------------------------------------------------------- pg_video
|
||||
; 256x256, 256 colours, 31.5 kHz, the two pages scrolled 384 apart, page 1 on
|
||||
; top. Every value here is tools/bench/crtc_mode.lua's, which derives them from
|
||||
; the dot clocks rather than recalling them; this is the same table in the place
|
||||
; a player would keep it.
|
||||
pg_video:
|
||||
movem.l d0-d1/a0-a1,-(sp)
|
||||
lea pg_crtc(pc),a0
|
||||
lea CRTC,a1
|
||||
moveq #0,d0
|
||||
.c: move.w (a0)+,d0
|
||||
bmi.s .cdone
|
||||
move.w (a0)+,d1
|
||||
move.w d1,0(a1,d0.w)
|
||||
bra.s .c
|
||||
.cdone:
|
||||
move.w #PG_R20D,CRTC_R20.l
|
||||
move.w #$0001,VC0.l ; 256 colours
|
||||
; PAGE 1 ON TOP. MEASURED (tools/bench/probe_page1.lua): $0000 puts page 0 on
|
||||
; top, its zeros then cover page 1 and the right half of the screen is black --
|
||||
; which was the first failure of the packed layout and does not look like a
|
||||
; priority bug from the outside.
|
||||
move.w #$0002,VC1.l
|
||||
move.w #$001F,VC2.l ; graphics on, all four pages
|
||||
move.b #15,CONTRAST.l ; the IPL leaves 14, i.e. 7% dark
|
||||
; The scroll. Both registers of each page, for the reason named at CR_P0X0.
|
||||
; move.w #0 and not clr.w: CLR on a 68000 READS the destination before writing
|
||||
; it, and a write-only register read back is a machine-dependent value this code
|
||||
; has no reason to depend on.
|
||||
move.w #0,CR_P0X0.l
|
||||
move.w #0,CR_P0Y0.l
|
||||
move.w #0,CR_P0X1.l
|
||||
move.w #0,CR_P0Y1.l
|
||||
move.w #384,CR_P1X0.l ; X = 384 = -128 mod 512
|
||||
move.w #0,CR_P1Y0.l
|
||||
move.w #384,CR_P1X1.l
|
||||
move.w #0,CR_P1Y1.l
|
||||
movem.l (sp)+,d0-d1/a0-a1
|
||||
rts
|
||||
|
||||
; register number * 2, value; -1 ends it
|
||||
pg_crtc:
|
||||
dc.w 0*2,45 ; H total: 46 chars = 368 dots = 31500 Hz
|
||||
dc.w 1*2,5 ; H sync end
|
||||
dc.w 2*2,10 ; H disp begin
|
||||
dc.w 3*2,42 ; H disp end -> 256 dots
|
||||
dc.w 4*2,567 ; V total: 568 lines -> 55.4577 Hz
|
||||
dc.w 5*2,5 ; V sync end
|
||||
dc.w 6*2,40 ; V disp begin
|
||||
dc.w 7*2,552 ; V disp end -> 512 lines, double-scanned
|
||||
dc.w 8*2,27 ; H sync adjust
|
||||
dc.w -1,0
|
||||
|
||||
; ---------------------------------------------------------------- pg_static
|
||||
; The half of the screen the container does not carry (FINDINGS 47.2). Two
|
||||
; regions, and they are static for two different reasons:
|
||||
;
|
||||
; words 128..511 of EVERY row. Page 1's storage at 384..511 sits under screen
|
||||
; columns 0..127 after the +384 scroll and must read 0 so the opaque page 0
|
||||
; shows through; the rest is off-screen storage. Zero, once.
|
||||
; the LETTERBOX rows, 0..31 and 224..255. The picture is 192 rows of a
|
||||
; 256-row screen. $FFFF is index 255 in both bytes, and 255 is black in
|
||||
; every frame's palette (vq.frame_palette reserves it), so these rows stay
|
||||
; correct across a per-frame palette without being rewritten.
|
||||
;
|
||||
; Written with the WRITE WINDOW OPEN, because that is the only mode in which a
|
||||
; word write lands both bytes -- outside it gvram_w takes `data & 0x00ff` and
|
||||
; the high byte is thrown away, which is the whole 2.0 B/pixel problem (46.5).
|
||||
pg_static:
|
||||
movem.l d0-d3/a0-a1,-(sp)
|
||||
move.w #PG_R20B,CRTC_R20.l
|
||||
lea GVRAM,a0
|
||||
move.w #SCRH-1,d0 ; y
|
||||
moveq #0,d3
|
||||
.row:
|
||||
movea.l a0,a1
|
||||
move.w d3,d1
|
||||
cmp.w #PG_YOFF,d1
|
||||
bcs.s .letter
|
||||
cmp.w #PG_YOFF+PG_H,d1
|
||||
bcc.s .letter
|
||||
lea PG_ROWB*2(a1),a1 ; a picture row: leave words 0..127 to
|
||||
move.w #SCRW-PG_ROWB-1,d1 ; the channel, zero 128..511
|
||||
bra.s .z
|
||||
.letter:
|
||||
move.w #PG_ROWB-1,d1 ; words 0..127 <- black
|
||||
.lb: move.w #PG_BLACK,(a1)+
|
||||
dbra d1,.lb
|
||||
move.w #SCRW-PG_ROWB-1,d1 ; ...and 128..511 <- 0, as everywhere
|
||||
.z: clr.w (a1)+
|
||||
dbra d1,.z
|
||||
lea PG_STRIDE(a0),a0
|
||||
addq.w #1,d3
|
||||
dbra d0,.row
|
||||
move.w #PG_R20D,CRTC_R20.l
|
||||
movem.l (sp)+,d0-d3/a0-a1
|
||||
rts
|
||||
|
||||
include "src/player/clock.i"
|
||||
; scsi.i includes src/player/dma.i itself, at its foot -- so it is NOT included
|
||||
; here. The DMAC constants and sc_in_dma come in with the transport that uses
|
||||
; them, which is also why dmagate.s includes only scsi.i.
|
||||
include "src/player/scsi.i"
|
||||
Reference in New Issue
Block a user