Two sessions, unrecorded until now, committed together because their edits share files and cannot be split cleanly after the fact. Session 28 (FINDINGS 60): the container is DLX5 -- every record sector-aligned, 120/120 starting on a boundary where 3/120 did, +0.48% on the wire and zero clocks -- and the ring's release rounds to RECALN so no pad is stranded. Two encoder levers measured and refused: `--spans all` buys +0.19 dB for +67% of the wire, and joint span/lam selection emits byte-identical containers because `lam` never leaves its floor on any of 120 frames. Session 29 (FINDINGS 61): the packed full-frame blit is 27.3% of a 12 fps frame, a channel fills GVRAM in buffer mode off the disc with the CPU halted, and it walks the 1,024 B line stride itself through array chaining. At the 9 clk/B dual-address floor the codec is 110.4% of a frame and a decoder-free packed literal player is 55.2%, at +4.89 dB -- 2.75 dB past a ceiling the codec's scene-wide palette cannot cross. Encoder work is parked; the codec is kept and not built on. check.sh is ALL GREEN before and after, plus one new stage that gates the ORDER of the measured paint costs rather than their values. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
151 lines
7.2 KiB
Python
151 lines
7.2 KiB
Python
"""A record is not a sector: what the mismatch costs, three ways (P4b, 58.3).
|
|
|
|
src/player/ring.i asks the transport for a RECORD -- a byte offset into the
|
|
scene's frame stream and a length, both 4-byte aligned because that is what
|
|
`move.l (a0)+` needs (28.3) and neither of them a multiple of 512. A SCSI
|
|
target answers in 512 B BLOCKS. On the gate container 117 of 120 records start
|
|
part way into a sector, so something has to reconcile the two, and the three
|
|
ways of doing it are not close.
|
|
|
|
WHY IT IS NOT AN IMPLEMENTATION DETAIL. The bytes on either side of a record in
|
|
the stream belong to OTHER records -- ones the decoder may still be reading --
|
|
and the block loop walks a0 with no bounds check at all (49.2). So a transport
|
|
that reads whole sectors straight into the ring does not waste 500 bytes, it
|
|
CORRUPTS the neighbours, and the symptom is wrong pixels rather than a fault.
|
|
|
|
A. WINDOWED PIO. Read the sectors the record lies in, store only the record.
|
|
src/player/scsi.i does this and it is what FINDINGS 58 measured. It costs
|
|
nothing in clocks -- the CPU is touching every byte anyway -- and it costs
|
|
the extra sectors on the wire. It CANNOT be done by a DMAC: a channel
|
|
writes a contiguous run to a contiguous address and cannot be told to drop
|
|
the first 300 bytes.
|
|
B. BOUNCE BUFFER. Let the DMAC write whole sectors somewhere else, then copy
|
|
the record into the ring. Works under DMA, and costs a copy of every
|
|
delivered byte -- which is precisely the cost `aligned` was chosen over
|
|
`split` to avoid (49.3, 19_ring_stream.py).
|
|
C. SECTOR-ALIGNED RECORDS. Pad each record up to 512 in the container
|
|
instead of up to 4. Costs bytes on the disc and in every delivery, and
|
|
nothing else at all; the transport becomes a whole-sector read into the
|
|
ring with no window and no copy. It is a CONTAINER change -- a re-encode
|
|
and a re-measurement of every constant fitted to the gate container, which
|
|
is the class of change ROADMAP already has bundled with P2's other half.
|
|
|
|
python3 tools/analysis/26_sector_align.py <in.dlx> [--ring KB]
|
|
|
|
No rate is taken and none is needed: every figure here is a fraction of the
|
|
delivered bytes or a count of clocks, and both are rate-free. What a given
|
|
delivery rate does with them is 15_bus_occupancy.py's question.
|
|
"""
|
|
import sys, os, argparse
|
|
sys.path.insert(0, "tools/encoder")
|
|
from dlx import DLX
|
|
|
|
SECTOR = 512
|
|
CPUHZ = 10_000_000
|
|
# 5.0 clocks/byte, and it is 19_ring_stream.py's constant rather than a new one:
|
|
# a 68000 `move.l (a0)+,(a1)+` moves 4 bytes in 20 clocks on a 16-bit bus. It
|
|
# is the OPTIMISTIC figure there and it is the optimistic figure here.
|
|
COPY_CLK_PER_BYTE = 5.0
|
|
# The windowed PIO loop in src/player/scsi.i, from the 68000's cycle table:
|
|
# 12 move.l #SC_PATIENCE,d3 patience reload
|
|
# 16 move.b SC_SSTS,d0 (xxx).L -> Dn
|
|
# 10 btst #0,d0
|
|
# 10 beq.s taken
|
|
# 20 move.b SC_DREG,(a1)+ (xxx).L -> (An)+
|
|
# 8 subq.l #1,d7
|
|
# 10 bne.s taken
|
|
# FINDINGS 58.2 measured 87.28 clocks per delivered byte against this loop's 86
|
|
# plus 1.15 for the dropped window bytes -- 0.2% apart, which is what says the
|
|
# cost is the instruction stream and not MAME's device model.
|
|
PIO_CLK_PER_BYTE = 86.0
|
|
|
|
ap = argparse.ArgumentParser()
|
|
ap.add_argument("container")
|
|
ap.add_argument("--ring", type=int, default=256, help="ring size in KB")
|
|
a = ap.parse_args()
|
|
|
|
d = DLX(a.container)
|
|
|
|
# The disc layout the 68000 walks: [u32 len][body], each record padded up to the
|
|
# container's own alignment -- 4 on DLX2/3/4, 512 on DLX5. Exactly
|
|
# tools/bench/prep_stream.py's, and it comes from the reader rather than from a
|
|
# second copy of the rule here, so pointing this tool at a DLX5 container asks
|
|
# it the RIGHT question: what does the mismatch still cost once the container
|
|
# has been changed to remove it? (The answer had better be nothing.)
|
|
off, recs = 0, []
|
|
for ln in d.record_lengths():
|
|
recs.append((off, ln))
|
|
off += ln
|
|
# The DENOMINATOR is the record bytes the decoder actually reads -- [u32 len]
|
|
# plus payload -- and NOT the padded length, because on a DLX5 container the
|
|
# padding IS the cost being measured. Scoring against the padded length would
|
|
# make an already-aligned container report +0.00% and look free.
|
|
payload = sum(4 + n for _, n in d.frames)
|
|
nfr = len(recs)
|
|
budget = CPUHZ / d.fps
|
|
|
|
print(f"{a.container}: {nfr} records, {payload:,} B, {d.fps} fps")
|
|
print(f" mean record {payload/nfr:,.0f} B; a {d.fps} fps frame is "
|
|
f"{budget:,.0f} clocks")
|
|
aligned0 = sum(1 for o, _ in recs if o % SECTOR == 0)
|
|
print(f" records that already start on a sector boundary: {aligned0}/{nfr}")
|
|
print()
|
|
|
|
# ---- A. windowed PIO: the sectors the record lies in, and only the record kept
|
|
wire_a = sum(((o % SECTOR) + ln + SECTOR - 1) // SECTOR for o, ln in recs) * SECTOR
|
|
drop_a = wire_a - payload
|
|
print("A. WINDOWED PIO (src/player/scsi.i, what FINDINGS 58 ran)")
|
|
print(f" wire {wire_a:,} B for {payload:,} B of record "
|
|
f"= +{100*drop_a/payload:.2f}%")
|
|
print(f" clocks {PIO_CLK_PER_BYTE:.0f}/B on EVERY byte off the FIFO, "
|
|
f"dropped ones included:")
|
|
print(f" {PIO_CLK_PER_BYTE*wire_a/nfr:,.0f} clk/frame "
|
|
f"= {100*PIO_CLK_PER_BYTE*wire_a/nfr/budget:.0f}% of the frame")
|
|
print( " and it does not survive the move to the DMAC at all: a channel "
|
|
"cannot drop bytes.")
|
|
print()
|
|
|
|
# ---- B. bounce buffer: DMA whole sectors elsewhere, copy the record in
|
|
print("B. BOUNCE BUFFER (whole sectors by DMA, then a copy)")
|
|
print(f" wire {wire_a:,} B, the same +{100*drop_a/payload:.2f}% -- the "
|
|
f"command is identical")
|
|
print(f" clocks {COPY_CLK_PER_BYTE:g}/B of copy on every DELIVERED byte, "
|
|
f"on top of whatever W the")
|
|
print(f" channel steals: {COPY_CLK_PER_BYTE*payload/nfr:,.0f} clk/frame "
|
|
f"= {100*COPY_CLK_PER_BYTE*payload/nfr/budget:.1f}% of the frame")
|
|
print( " which is the cost `aligned` was chosen over `split` to avoid "
|
|
"(49.3), arriving")
|
|
print( " by a different door and on every byte instead of on a wrap.")
|
|
print()
|
|
|
|
# ---- C. sector-aligned records in the container
|
|
cur, pad = 0, 0
|
|
for _, ln in recs:
|
|
if cur % SECTOR:
|
|
pad += SECTOR - (cur % SECTOR)
|
|
cur += SECTOR - (cur % SECTOR)
|
|
cur += ln
|
|
print("C. SECTOR-ALIGNED RECORDS (a container change; a re-encode)"
|
|
+ (" -- THIS CONTAINER ALREADY IS ONE" if d.sector_aligned else ""))
|
|
print(f" wire {cur:,} B for {payload:,} B of record = "
|
|
f"+{100*(cur-payload)/payload:.2f}%")
|
|
print( " clocks ZERO: the read is a whole-sector read straight into the "
|
|
"ring, no window,")
|
|
print( " no copy, and the DMAC can do it.")
|
|
print()
|
|
|
|
ringsz = a.ring * 1024
|
|
print(f" VERDICT, in the currency this project prices delivery in. C is "
|
|
f"cheaper on the wire")
|
|
print(f" than A and B by {100*(wire_a-cur)/payload:.2f} points of the payload "
|
|
f"({wire_a-cur:,} B on this scene),")
|
|
print(f" and it is the only one of the three a DMA channel can run without a "
|
|
f"copy. What it")
|
|
print(f" costs is a container revision and the re-measurement that comes with "
|
|
f"one.")
|
|
maxrec = max(ln for _, ln in recs)
|
|
maxpad = maxrec + (-maxrec) % SECTOR
|
|
print(f" It also grows the largest record from {maxrec:,} to {maxpad:,} B, "
|
|
f"which a {a.ring} KB")
|
|
print(f" ring still holds {ringsz//maxpad} times over.")
|