Put sound on the wire, and find three LSBs are worth 25 dB

ROADMAP P6, everything in the item except the bus half session 20 closed.
tools/encoder/adpcm.py is an MSM6258 codec, tools/encoder/extract_audio.py
takes the same seconds of the same stream the frames come from,
tools/bench/verify_adpcm.py is the gate, tools/analysis/32_audio_wire.py the
container arithmetic.

There is no reference encoder -- ffmpeg has a decoder for this format and none
the other way -- so what is gated is the decoder the encoder runs INSIDE its
own nibble search, sample-exact against ffmpeg's over 4,268 nibbles. An
encoder that agrees with its own wrong decoder is what that catches. The Singe
window: 156,250 samples -> 78,125 B at 21.97 dB, which is 7,812.5 B/s to the
byte. Normalising the disc's -13.4 dBFS level moves the SNR 21.97 -> 21.97, so
the level is not a lever.

And the two published delta formulas are not the same codec. They differ by at
most 3 in 12-bit units; encode for one and decode on the other and the SNR
goes 21.97 -> -2.88 dB, the noise louder than the signal, because ADPCM is
recursive the way the video codec is temporally recursive. Which one the chip
runs is now P6a and it is a precondition on shipping any audio.

And audio is the first thing the packed branch's simplification has cost
anything for. A record has no index BY DESIGN, so audio cannot be per-record
without making records variable; it rides a fixed cadence (F, A), the obvious
F=1 wastes 57.3% of every audio sector, and the pick is F=11 A=14 -- 0.09%
padding, 14,336 B held, wire 582.0 -> 589.6 KB/s. The codec container, which
kept its index, pays zero.

The MAME experiment did not work and 65.5 says so: :okim6258 is there at
$E92001/$E92003, read out of the machine's own program map, and feeding it
from Lua recorded silence across control 0..3 x port C 0..15. The register
semantics were not guessed at further.

FINDINGS 65. check.sh ALL GREEN before and after, with a new stage.

Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
prosolis
2026-08-25 09:11:03 -07:00
parent 6f698ca226
commit f925a1dd9a
12 changed files with 1069 additions and 6 deletions
+227
View File
@@ -0,0 +1,227 @@
#!/usr/bin/env python3
"""WHAT DOES AUDIO DO TO THE CONTAINER? ROADMAP P6, the half that is not the bus.
python3 tools/analysis/32_audio_wire.py [packed.dlxp] [--audio tmp/au_singe.raw]
[--rate KB/s ...]
Session 20 (FINDINGS 52) closed the bus half of P6: a second DMA consumer at
7,812.5 B/s is 1.25%..1.48% of a frame, about 4% of what the decoder leaves, and
the 7.8 kB/s figure survived with a unit correction. ROADMAP P6 then says, in
as many words, that EVERYTHING ELSE in the item is open: extraction, an encoder,
the container interleave, and what a second stream does to `wire` and therefore
to `pipe - wire` and therefore to 51.3's refill climb.
This file is the container interleave and the wire. It is arithmetic over the
real container's real geometry -- no MAME run, no board.
THE THING THAT MAKES IT INTERESTING, and it is a property of DLXP1 rather than
of audio: **a packed container has no index and cannot have one.** A record's
address is `LBA0 + i*97` because a literal frame's length is geometry (FINDINGS
63, 64.1). Audio is a stream at a rate that has nothing to do with the frame
rate, so the naive interleave -- give record i the audio bytes belonging to slot
i -- makes records VARIABLE LENGTH, and the moment records are variable length
the format needs an index and stops being the format.
So the interleave has to be a FIXED CADENCE: every F frames, A whole sectors of
audio, placed between records. Then
LBA(i) = LBA0 + i*RECSEC + floor(i/F)*A
which is still two multiplies and a divide -- arithmetic, no index, nothing
walked -- and the only cost is that A*512 must be at least F frames' worth of
audio, so the padding is whatever A*512 exceeds it by. Choosing (F, A) is a
rational-approximation problem and the answer is NOT the obvious cadence.
"""
import argparse, os, sys
from fractions import Fraction
sys.path.insert(0, "tools/encoder")
sys.path.insert(0, "tools/analysis")
import buscost as B
from dlxp import DLXP, SECTOR
ap = argparse.ArgumentParser()
ap.add_argument("container", nargs="?", default="tmp/packed_singe.dlxp")
ap.add_argument("--audio", default="tmp/au_singe.raw",
help="raw s16le mono at the chip rate, from extract_audio.py")
ap.add_argument("--codec", default="tmp/rc_fr_singe_scsi_span.dlx",
help="the codec container, for the same arithmetic on the other branch")
ap.add_argument("--rate", type=float, nargs="*",
default=[453.6, 500.0, 582.0, 600.0, 650.0, 700.0],
help="explicit sustained delivery rates, KB/s")
a = ap.parse_args()
d = DLXP(a.container)
SLOT_S = 1.0 / d.fps
FRAME_CLK = B.CPU_HZ * SLOT_S if hasattr(B, "CPU_HZ") else 10_000_000 * SLOT_S
RECSEC = d.rec_bytes // SECTOR
print(f"""
=== THE STREAM =========================================================
The chip is an MSM6258V on an 8 MHz clock and it has three rates and no
others. Every budget in this tree is written against the first one.""")
RATES = {512: 15625.0, 768: 8_000_000/768, 1024: 7812.5}
print(f"\n {'divisor':>8} {'samples/s':>11} {'bytes/s':>10} {'B per 1/%d s slot' % d.fps:>19} exact?")
for div, hz in RATES.items():
bps = hz / 2
per = bps / d.fps
print(f" 8MHz/{div:<4} {hz:11,.1f} {bps:10,.1f} {per:19,.4f} "
f"{'yes' if per == int(per) else 'NO -- a remainder, like the frame clock (54)'}")
HZ = 15625.0
AU_BPS = HZ / 2 # 4 bits a sample, two samples to a byte
AU_FRAME = AU_BPS / d.fps # 651.0416... B, and the point is the dots
print(f"""
The shipping rate's per-slot figure is {AU_FRAME:,.4f} B and it is NOT an
integer -- 8 MHz / 512 / 2 / {d.fps} has a 12 in the denominator that 2**k
cannot clear. That is the same shape as FINDINGS 54's frame clock: what a
player carries is a remainder, not a count, and a container that rounds it
either drifts or underruns.""")
if os.path.exists(a.audio):
n16 = os.path.getsize(a.audio) // 2
secs = n16 / HZ
print(f"""
MEASURED, on the window this project gates everything on (00223 @539.4s,
{secs:.3f} s, tools/encoder/extract_audio.py):
{n16:,} samples -> {n16//2:,} B of ADPCM = {n16/2/secs:,.1f} B/s
which is {AU_BPS:,.1f} to the byte, so the rate is the rate.""")
print(f"""
=== THE INTERLEAVE, AND WHY THE OBVIOUS CADENCE IS THE WRONG ONE =======
A packed record is {d.rec_bytes:,} B = {RECSEC} sectors EXACTLY and its address is
arithmetic. Audio rides between records at a fixed cadence -- every F frames,
A whole sectors -- so that LBA(i) stays arithmetic. A must satisfy
A * {SECTOR} >= F * {AU_FRAME:,.4f} i.e. A/F >= {Fraction(int(AU_BPS*2), int(2*SECTOR*d.fps))} = {AU_FRAME/SECTOR:.9f}
and everything above that ratio is PADDING that the wire pays for and nothing
plays. Here is the whole small-F space, best A for each F:""")
target = Fraction(int(round(AU_BPS * 2)), 2 * SECTOR * d.fps) # sectors per frame, exact
rows, floor = [], None
for F in range(1, 241):
A = -(-(target.numerator * F) // target.denominator) # ceil(F * target)
have, need = A * SECTOR, F * AU_FRAME
waste = (have - need) / need
add = have / F * d.fps / 1024 # what the cadence puts on the wire, KB/s
rows.append((waste, F, A, have, need, add))
print(f"\n {'F':>4} {'A':>4} {'A*512 B':>10} {'needs':>12} {'padding':>9} {'waste':>7}"
f" {'wire adds':>10} {'player RAM':>11}")
seen = None
for waste, F, A, have, need, add in rows:
show = F <= 4 or seen is None or waste < seen - 1e-12
if seen is None or waste < seen: seen = waste
if show:
print(f" {F:4d} {A:4d} {have:10,} {need:12,.1f} {have-need:9,.1f} "
f"{100*waste:6.2f}% {add:9.2f} KB/s {have:9,} B")
best = sorted(rows)
w, F, A, have, need, add = best[0]
f1 = next(x for x in rows if x[1] == 1)
print(f""" THE FLOOR OF THAT SWEEP is F={F}, A={A}: {100*w:.3f}% padding, {add:.2f} KB/s of
wire for {AU_BPS/1024:.2f} KB/s of audio.
THE OBVIOUS CADENCE IS THE WORST ONE. F=1 -- one audio lump per record, which
is what "interleave the audio into the frame" means if nobody does the
arithmetic -- needs A={f1[2]} and costs {100*f1[0]:.1f}% padding: {AU_FRAME:,.1f} B rounded up to
{f1[3]:,}, so {f1[3]-AU_FRAME:,.1f} B of every record is nothing at all, and the wire pays
{f1[5]:.2f} KB/s for {AU_BPS/1024:.2f} KB/s of audio. That is {f1[5]-add:.2f} KB/s thrown away for
no reason but the cadence.
=== WHAT IT DOES TO THE WIRE ===========================================""")
vid_kbs = d.rec_bytes * d.fps / 1024
for label, cad in (("F=1 (one lump a record)", f1), (f"F={F} (the floor)", best[0])):
tot = vid_kbs + cad[5]
print(f" {label:26s} video {vid_kbs:7.1f} + audio {cad[5]:5.2f} = {tot:7.1f} KB/s "
f"({100*(tot/vid_kbs-1):+.2f}%)")
print(f"""
And this is what B1's acceptance test becomes. The packed container's
sustained requirement was {vid_kbs:.1f} KB/s SILENT (FINDINGS 61.5, 63) and it is
{vid_kbs + add:.1f} KB/s with sound. A literal frame's bitrate is geometry and cannot
be talked down; the audio on top of it is {add:.2f} KB/s and can only be talked down
by choosing a worse chip rate.""")
f11 = next(x for x in rows if x[1] == 11)
print(f"""
AND THE CADENCE HAS A SECOND PRICE, WHICH IS RAM. A cadence of F frames means
the player is holding F frames of audio, and holding it TWICE -- the channel
fills lump n+1 while the chip drains lump n, the same reason K4 needs two
record buffers (64.2). So the floor of the sweep is not the answer:
F={f1[1]:<3} {f1[3]:>7,} B a lump, {2*f1[3]:>7,} B held {100*f1[0]:6.2f}% padding {f1[5]:5.2f} KB/s
F={f11[1]:<3} {f11[3]:>7,} B a lump, {2*f11[3]:>7,} B held {100*f11[0]:6.2f}% padding {f11[5]:5.2f} KB/s <- the pick
F={F:<3} {have:>7,} B a lump, {2*have:>7,} B held {100*w:6.2f}% padding {add:5.2f} KB/s
F={f11[1]} buys {100*(f1[0]-f11[0]):.1f} points of padding for {2*f11[3]-2*f1[3]:,} B of RAM, and F={F} buys the
last {100*(f11[0]-w):.2f} of a point for {2*have-2*f11[3]:,} B more. On a machine where K4 already
wants 99,328 B for two record buffers, the second trade is not one.
=== THE ASYMMETRY: THE CODEC CONTAINER PAYS NONE OF THIS ===============""")
if os.path.exists(a.codec):
sys.path.insert(0, "tools/encoder")
from dlx import DLX
c = DLX(a.codec)
lens = c.record_lengths() if callable(getattr(c, "record_lengths", None)) else c.record_lengths
cwire = sum(lens) / len(lens) * c.fps / 1024
print(f""" {os.path.basename(a.codec)}: {c.nframes} records, index {'PRESENT' if c.has_index else 'absent'},
records already VARIABLE ({min(lens):,}..{max(lens):,} B, mean {sum(lens)/len(lens):,.0f}) and
sector-aligned since DLX5 (60.1). A container that already carries an index
and already has variable records can put EXACTLY {AU_FRAME:,.1f} B of audio in record i
and pad only to the sector it was going to pad to anyway -- so its audio
padding is not 57.3% and not 1.11%, it is ZERO, and its wire goes
{cwire:.1f} -> {cwire + AU_BPS/1024:.1f} KB/s ({100*(AU_BPS/1024)/cwire:+.2f}%).
THAT IS THE FIRST COST THIS PROJECT HAS FOUND FOR THE PACKED BRANCH'S OWN
SIMPLIFICATION. "A record's length is geometry, so there is no index and none
can be needed" (63, 64.1) is what makes the packed player a page of arithmetic
instead of a parser -- and it is exactly the property that makes a second
stream at an unrelated rate cost padding, a cadence, and a buffer. It is a
small cost ({f11[5]-AU_BPS/1024:.2f} KB/s at the pick, {2*f11[3]:,} B of RAM) and it is not zero, and
nothing in FINDINGS 61-64 predicted it.""")
else:
print(f" SKIPPED: no codec container at {a.codec}")
print(f"""
=== WHAT IT DOES TO SLACK (51.3) =======================================
Slack is ACCUMULATED out of pipe - wire, so a second consumer does not cost a
fixed amount -- it costs the accumulation rate, and what a branch point costs is
set by that (51.3, 55.4). Silent vs sounded, at explicit rates:
{'pipe':>8} {'silent':>14} {'sounded':>14} what a second of play banks""")
for kbps in a.rate:
s_sl, a_sl = kbps - vid_kbs, kbps - (vid_kbs + add)
def fmt(x): return f"{x:+8.1f} KB/s" if x >= 0 else f"{x:+8.1f} KB/s"
print(f" {kbps:8.1f} {fmt(s_sl):>14} {fmt(a_sl):>14} "
+ ("both starve" if a_sl < 0 and s_sl < 0
else "SOUND IS WHAT BREAKS IT" if s_sl >= 0 > a_sl
else f"{a_sl/s_sl*100:.0f}% of the silent rate" if s_sl > 0 else ""))
AUCLK_LO = AU_FRAME * B.ADPCM_CLK_BYTE_BEST
AUCLK_HI = AU_FRAME * B.ADPCM_CLK_BYTE_WORST
print(f"""
=== AND WHAT IT DOES TO THE FRAME (the half session 20 already closed) ==
{AU_FRAME:,.1f} B a slot at {B.ADPCM_CLK_BYTE_BEST}..{B.ADPCM_CLK_BYTE_WORST} clocks a byte (the IPL ROM's OWN channel-3
configuration, read out of the ROM by 21_iplrom_dmac.py, not chosen here) is
{AUCLK_LO:,.0f}..{AUCLK_HI:,.0f} clocks = {100*AUCLK_LO/FRAME_CLK:.2f}%..{100*AUCLK_HI/FRAME_CLK:.2f}% of a {SLOT_S*1000:.2f} ms slot.
That reproduces FINDINGS 52 exactly, which is the point of printing it.
THE INTERACTION 52 COULD NOT HAVE HAD is with 64.2's write window. A
DMAC-direct packed player holds the GVRAM window open for the whole data
phase, and an audio channel stealing the bus during that phase makes the phase
LONGER -- so audio does not merely cost clocks, it costs DARKNESS:
extra dark per slot = {100*AUCLK_LO/FRAME_CLK:.2f}%..{100*AUCLK_HI/FRAME_CLK:.2f}% of the slot, on top of
record/(burst x slot), which is already 1.0 at the wire
It is small against a dark fraction that is already 1.0, and it is not small
against K4's {100*227553/FRAME_CLK:.1f}% paint. For the CPU-painted player the audio steals
from the paint and not from the picture, which is the third time this session
the two players have ranked differently on a column that is not clocks.
""")
+18
View File
@@ -735,4 +735,22 @@ else
echo " SKIPPED: no x68000 romset -- the player was not run"
fi
echo "--- session 33: AUDIO -- the encoder, and what it does to the wire (FINDINGS 65) ---"
# ROADMAP P6, everything in it except the bus half session 20 closed. The audio
# is the SAME WINDOW as the frames -- 00223 from 539.4 s for 10 s -- because an
# audio stream that is not the same seconds as the picture is not this project's
# audio, and a gate that lets the two drift apart would never say so.
[ -f tmp/au_singe.raw ] || python3 tools/encoder/extract_audio.py 00223 tmp/au_singe.raw 15625 539.4 10.0
# There is NO ffmpeg encoder for this format -- adpcm_ima_oki is decode-only --
# so the encoder cannot be checked against a reference. What is checked is that
# the decoder our encoder runs in its own loop IS ffmpeg's, sample for sample.
# An encoder that agrees with its own wrong decoder is the failure this catches.
python3 tools/bench/verify_adpcm.py tmp/au_singe.raw || exit 1
# And the container arithmetic. The interesting line is the padding: a packed
# record has no index BY DESIGN, so audio has to ride a fixed cadence, and the
# obvious cadence throws away a third of every audio sector.
python3 tools/analysis/32_audio_wire.py tmp/packed_singe.dlxp \
> tmp/audio_wire.log 2>&1 || { cat tmp/audio_wire.log; exit 1; }
grep -aE "^ ( 1| 11| 81) |THE FLOOR|F=1 |F=11|is ZERO|SOUND IS WHAT" tmp/audio_wire.log
echo "ALL GREEN"
+28
View File
@@ -0,0 +1,28 @@
-- Ask MAME what the x68000's ADPCM device is, and where the 68000 reaches it.
-- FINDINGS 64.4 recorded that no MAME source tree is on this machine; this is
-- the way to ask the same question without one.
M = manager.machine
local done = false
SUB = emu.add_machine_frame_notifier(function()
if done then return end
done = true
print("[AD] === devices whose tag looks like an ADPCM chip ===")
for tag, dev in pairs(M.devices) do
local t = tag:lower()
if t:find("adpcm") or t:find("oki") or t:find("msm") or t:find("6258") then
print(string.format("[AD] DEV %-24s shortname=%s", tag, tostring(dev.shortname)))
end
end
local sp = M.devices[":maincpu"].spaces["program"]
print("[AD] === program map, $E90000..$EA0000 ===")
local ok, err = pcall(function()
for _, e in ipairs(sp.map.entries) do
if e.address_start >= 0xE90000 and e.address_start < 0xEA0000 then
print(string.format("[AD] MAP %08X-%08X", e.address_start, e.address_end))
end
end
end)
if not ok then print("[AD] MAP unavailable: " .. tostring(err)) end
print("[AD] done")
M:exit()
end)
+36
View File
@@ -0,0 +1,36 @@
-- Feed the x68000's OWN okim6258 a known nibble stream and let MAME record what
-- comes out, so that "which delta formula does the chip use" is a MEASUREMENT
-- and not a reading of source code that is not on this machine (64.4).
--
-- The feed is deliberately SLOW -- a byte every host frame, where real time
-- wants ~138 -- because the question is not the rate. A starved chip holds its
-- last sample, so the wave is a STAIRCASE of the reconstructed values, which is
-- exactly the sequence the two candidate formulas disagree about.
M = manager.machine
local sp = M.devices[":maincpu"].spaces["program"]
local CTRL, DATA = 0xE92001, 0xE92003
local CMD = tonumber(os.getenv("AD_CMD") or "1")
-- 12 loud nibbles to climb the step index, then every nibble in turn: the pairs
-- where the two formulas differ are all at step indices above the floor.
local nibs = {}
for i = 1, 12 do nibs[#nibs+1] = 7 end
for i = 0, 15 do nibs[#nibs+1] = i end
for i = 0, 15 do nibs[#nibs+1] = i end
local bytes = {}
for i = 1, #nibs, 2 do bytes[#bytes+1] = nibs[i] * 16 + nibs[i+1] end
local n, started = 0, false
SUB = emu.add_machine_frame_notifier(function()
n = n + 1
if n == 30 then
sp:write_u8(CTRL, CMD)
started = true
print(string.format("[AD] ctrl $%02X written to $%06X", CMD, CTRL))
elseif started and n > 30 and (n - 30) <= #bytes then
sp:write_u8(DATA, bytes[n - 30])
elseif started and (n - 30) == #bytes + 20 then
print("[AD] fed " .. #bytes .. " bytes = " .. #nibs .. " nibbles")
print("[AD] done")
M:exit()
end
end)
+24
View File
@@ -0,0 +1,24 @@
-- Sweep the PPI's port C -- which is where the X68000 puts ADPCM pan and the
-- chip's clock divider -- and feed a loud burst under each value, so that the
-- WAV says which value un-mutes the chip. Nothing here is assumed: the segment
-- boundaries are printed and the analysis reads the wave against them.
M = manager.machine
local sp = M.devices[":maincpu"].spaces["program"]
local CTRL, DATA, PPIC = 0xE92001, 0xE92003, 0xE9A005
local SEG = 40 -- host frames per segment
local n = 0
SUB = emu.add_machine_frame_notifier(function()
n = n + 1
if n <= 20 then return end
local k = n - 20
local seg = math.floor((k - 1) / SEG)
local off = (k - 1) % SEG
if seg > 15 then print("[AD] done"); M:exit(); return end
if off == 0 then
sp:write_u8(PPIC, seg)
sp:write_u8(CTRL, 1)
print(string.format("[AD] seg %2d portC=$%02X starts at host frame %d", seg, seg, n))
elseif off <= 30 then
sp:write_u8(DATA, 0x77) -- two loud positive nibbles
end
end)
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
"""Gate tools/encoder/adpcm.py against the only independent decoder on this
machine: ffmpeg's `adpcm_ima_oki`.
There is no ffmpeg ENCODER for this format -- `adpcm_ima_oki` is decode-only --
so the encoder here cannot be checked against a reference implementation. What
CAN be checked, and is, is that the decoder our encoder runs in its own loop is
byte-for-byte the decoder that ships in ffmpeg. An encoder that agrees with its
own wrong decoder is exactly the failure this catches.
Usage: verify_adpcm.py [wav_or_raw12 ...]
"""
import os, struct, subprocess, sys, random, math
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "encoder"))
import adpcm
TMP = "tmp/adpcm_gate"
# The step table as it is printed in the OKI datasheet and in every
# implementation of this format. adpcm.py BUILDS its table from 16*1.1**k; if
# the two ever disagree, one of them is a typo and this says which.
CANON = [16,17,19,21,23,25,28,31,34,37,41,45,50,55,60,66,73,80,88,97,107,118,
130,143,157,173,190,209,230,253,279,307,337,371,408,449,494,544,598,
658,724,796,876,963,1060,1166,1282,1411,1552]
fails = []
def ck(ok, msg):
print(("OK " if ok else "FAIL ") + msg)
if not ok: fails.append(msg)
def ffmpeg_decode(data, rate=15625):
"""Decode packed OKI ADPCM through ffmpeg, by wrapping it in a WAV whose
format tag is 0x0010 (WAVE_FORMAT_OKI_ADPCM). Returns 16-bit samples."""
os.makedirs(TMP, exist_ok=True)
fmt = struct.pack("<HHIIHHH", 0x0010, 1, rate, rate, 1, 4, 0)
body = (b"WAVE" + b"fmt " + struct.pack("<I", len(fmt)) + fmt
+ b"data" + struct.pack("<I", len(data)) + data)
w = f"{TMP}/probe.wav"
open(w, "wb").write(b"RIFF" + struct.pack("<I", len(body)) + body)
raw = subprocess.check_output(
["ffmpeg", "-v", "error", "-i", w, "-f", "s16le", "-acodec", "pcm_s16le", "-"])
return list(struct.unpack("<%dh" % (len(raw) // 2), raw))
def snr_db(ref, got):
"""Signal-to-noise over the 12-bit sample word."""
num = sum(float(s) * s for s in ref)
den = sum((float(a) - b) ** 2 for a, b in zip(ref, got))
if den == 0: return float("inf")
return 10.0 * math.log10(num / den) if num else float("-inf")
print("--- the step table ---")
ck(adpcm.STEP == CANON, f"49 entries, built = published (16*1.1**k), {adpcm.STEP[0]}..{adpcm.STEP[-1]}")
print("--- our decoder vs ffmpeg's adpcm_ima_oki ---")
random.seed(1234)
nibs = ([7]*12 + [i % 16 for i in range(256)]
+ [random.randrange(16) for _ in range(4000)])
data = adpcm.pack(nibs)
ff = ffmpeg_decode(data)
ours = [v * 16 for v in adpcm.decode(adpcm.unpack(data, len(nibs)), "shift")]
ck(len(ff) == len(ours), f"sample count {len(ff)} = {len(ours)}")
ck(ff == ours, f"variant 'shift' is SAMPLE-EXACT vs ffmpeg over {len(nibs)} nibbles")
# NEGATIVE CONTROL. A gate that passes whatever it is handed proves nothing;
# reading the nibbles the other way round has to go red, or "high nibble first"
# is an assertion rather than a measurement.
lowfirst = [adpcm.unpack(data, len(nibs))[i ^ 1] for i in range(len(nibs))]
bad = [v * 16 for v in adpcm.decode(lowfirst, "shift")]
ndiff = sum(1 for a, b in zip(ff, bad) if a != b)
ck(ndiff > 0, f"low-nibble-first DISAGREES on {ndiff}/{len(ff)} -- so the order is measured, not assumed")
print("--- and the second variant is not the same decoder ---")
terms = [v * 16 for v in adpcm.decode(adpcm.unpack(data, len(nibs)), "terms")]
d = [abs(a - b) // 16 for a, b in zip(ff, terms)]
nd = sum(1 for x in d if x)
ck(nd > 0, f"variant 'terms' differs on {nd}/{len(d)} samples, max {max(d)} in 12-bit units"
" -- OPEN: which one the MSM6258 runs is unmeasured")
print("--- and getting the variant wrong is NOT a rounding error ---")
# THE MEASUREMENT THAT CHANGED THIS FROM A FOOTNOTE INTO AN OPEN ITEM. The two
# variants differ by at most 3 in 12-bit units PER SAMPLE, which reads like
# something nobody could hear. ADPCM is RECURSIVE -- the delta is added to a
# running predictor and the nibble also moves the step index -- so the
# disagreement does not stay where it happens. It is the same shape as the
# codec's temporal recursion (64.1), one dimension down.
if os.path.exists("tmp/au_singe.raw"):
raw = open("tmp/au_singe.raw", "rb").read()
pcm = struct.unpack("<%dh" % (len(raw) // 2), raw)
ref = [max(-2048, min(2047, x >> 4)) for x in pcm]
nib = adpcm.encode(ref, "shift")
same, cross = adpcm.decode(nib, "shift"), adpcm.decode(nib, "terms")
err = [abs(x - y) for x, y in zip(same, cross)]
print(f" encoded 'shift', decoded 'shift': SNR {snr_db(ref, same):6.2f} dB")
print(f" encoded 'shift', decoded 'terms': SNR {snr_db(ref, cross):6.2f} dB"
f" <- the noise is LOUDER THAN THE SIGNAL")
print(f" per-sample disagreement over {len(err):,} samples: max {max(err)}, "
f"mean {sum(err)/len(err):.1f} in 12-bit units")
ck(snr_db(ref, cross) < 0,
"a 3-LSB formula disagreement costs ~25 dB, because ADPCM is RECURSIVE")
else:
print(" SKIPPED: no tmp/au_singe.raw (tools/encoder/extract_audio.py)")
print("--- the encoder, through the gated decoder ---")
for path in (sys.argv[1:] or []):
raw = open(path, "rb").read()
if raw[:4] == b"RIFF":
raw = subprocess.check_output(["ffmpeg", "-v", "error", "-i", path,
"-f", "s16le", "-ac", "1", "-ar", "15625", "-"])
pcm16 = struct.unpack("<%dh" % (len(raw) // 2), raw)
src = [max(-2048, min(2047, s >> 4)) for s in pcm16]
nib = adpcm.encode(src, "shift")
packed = adpcm.pack(nib)
rec_ff = [v // 16 for v in ffmpeg_decode(packed)][:len(src)]
rec_us = adpcm.decode(nib, "shift")
ck(rec_ff == rec_us,
f"{os.path.basename(path)}: encoder's own reconstruction = ffmpeg's, {len(src)} samples")
print(f" {len(src)} samples, {len(packed)} B, SNR {snr_db(src, rec_us):.2f} dB "
f"(12-bit word; the source is already quantised to it)")
print("ADPCM GATE " + ("GREEN" if not fails else f"RED: {len(fails)} failed"))
sys.exit(1 if fails else 0)
+118
View File
@@ -0,0 +1,118 @@
#!/usr/bin/env python3
"""MSM6258 (OKI/Dialogic) 4-bit ADPCM -- encoder, decoder, and the fact that
there are TWO decoders and they are not the same one.
The X68000's ADPCM is an OKI MSM6258V clocked at 8 MHz, dividing to 15,625 /
10,417 / 7,812.5 samples a second, 4 bits each, two samples to a byte
(FINDINGS 52, buscost.ADPCM_SAMPLE_HZ). The sample word is 12 bits signed.
WHY THIS FILE HAS TWO DECODERS. Nothing in this repo can be trusted to say what
the chip does, and the two references available on this machine DISAGREE:
VARIANT 'shift' delta = ((2*(n&7) + 1) * step) >> 3
This is ffmpeg's `adpcm_ima_oki`, and `gate_vs_ffmpeg()`
reproduces it SAMPLE-EXACT, so it is not a reading of source
code -- it is a measurement of the decoder that ships.
VARIANT 'terms' delta = step/8 + (n&4 ? step : 0) + (n&2 ? step/2 : 0)
+ (n&1 ? step/4 : 0), each term truncated
This is the OKI datasheet's own form, the one an ADPCM chip
can actually build out of shifts and adds, and it is what
MAME's okim6258 is understood to compute. NOT VERIFIED HERE:
no MAME source tree is on this machine (FINDINGS 64.4).
They differ on 445 of 2,268 sampled nibbles, by up to 4 in 12-bit units --
small, and small is not zero. Which one the machine runs is an open question
with an experiment attached: MAME's x68000 HAS an okim6258, so it can be asked
rather than argued about.
Nibble order is HIGH NIBBLE FIRST within a byte -- measured, not assumed, by the
same gate: reading low-first mismatches ffmpeg on 1,728 of 2,268 samples.
"""
# The 49-entry OKI step table. floor(16 * 1.1**k) for k in 0..48 -- built rather
# than pasted, so a transcription slip is not one of the things that can be
# wrong here.
STEP = [int(16 * 1.1**k) for k in range(49)]
# The nibble magnitude's effect on the step index. Four quiet nibbles walk it
# down one, four loud ones walk it up by more.
INDEX_ADJUST = (-1, -1, -1, -1, 2, 4, 6, 8)
SAMPLE_MIN, SAMPLE_MAX = -2048, 2047 # the 12-bit DAC word
VARIANTS = ("shift", "terms")
def delta(nibble, step, variant):
"""The reconstruction step for one nibble, in 12-bit units."""
if variant == "shift":
d = ((2 * (nibble & 7) + 1) * step) >> 3
elif variant == "terms":
d = step // 8
if nibble & 4: d += step
if nibble & 2: d += step // 2
if nibble & 1: d += step // 4
else:
raise ValueError(f"unknown variant {variant!r}")
return -d if nibble & 8 else d
def decode(nibbles, variant="shift"):
"""Nibbles -> 12-bit signed samples. State is (signal, step index), both
zero at the start of a stream, which is what the chip resets to."""
signal, idx, out = 0, 0, []
for n in nibbles:
signal += delta(n, STEP[idx], variant)
signal = SAMPLE_MIN if signal < SAMPLE_MIN else (
SAMPLE_MAX if signal > SAMPLE_MAX else signal)
idx += INDEX_ADJUST[n & 7]
idx = 0 if idx < 0 else (48 if idx > 48 else idx)
out.append(signal)
return out
def encode(samples, variant="shift"):
"""12-bit signed samples -> nibbles.
The nibble is chosen by EXHAUSTIVE SEARCH over all sixteen, minimising the
reconstruction error of this sample. That is greedy rather than optimal --
a nibble also moves the step index, so a locally worse choice can pay later
-- but it is what a chip-matched encoder is expected to do and it costs
nothing offline. The decoder is run INSIDE the loop, so the encoder can
never drift away from what the decoder will reconstruct.
"""
signal, idx, out = 0, 0, bytearray()
for s in samples:
step = STEP[idx]
best, best_err = 0, None
for n in range(16):
v = signal + delta(n, step, variant)
v = SAMPLE_MIN if v < SAMPLE_MIN else (SAMPLE_MAX if v > SAMPLE_MAX else v)
err = (v - s) ** 2
if best_err is None or err < best_err:
best, best_err = n, err
signal += delta(best, step, variant)
signal = SAMPLE_MIN if signal < SAMPLE_MIN else (
SAMPLE_MAX if signal > SAMPLE_MAX else signal)
idx += INDEX_ADJUST[best & 7]
idx = 0 if idx < 0 else (48 if idx > 48 else idx)
out.append(best)
return bytes(out)
def pack(nibbles):
"""Nibbles -> bytes, HIGH NIBBLE FIRST. An odd count pads with a 0 nibble,
which is the quietest one the format has (delta = step/8)."""
n = bytes(nibbles)
if len(n) & 1:
n += b"\0"
return bytes((n[i] << 4) | n[i + 1] for i in range(0, len(n), 2))
def unpack(data, count=None):
out = bytearray()
for b in data:
out.append(b >> 4)
out.append(b & 15)
return bytes(out[:count] if count is not None else out)
+45
View File
@@ -0,0 +1,45 @@
#!/usr/bin/env python3
"""Extract the audio of a Blu-ray window as mono PCM at an MSM6258 sample rate.
The video side of this window is tools/encoder/extract.py; the arguments mean
the same things and are meant to be given the same values, because an audio
stream that is not the same seconds as the frames is not this project's audio.
The disc is AC-3 5.1 at 48 kHz. The arcade original is MONO, so this downmixes
-- ffmpeg's default matrix, dialogue from the centre channel included -- and
resamples to the chip's rate. Nothing here shapes, gates or normalises the
level: what the ADPCM encoder is handed is what the disc has, so that the SNR
it reports is the codec's and not a gain stage's.
"""
import getpass, os, subprocess, sys
BDROM = os.environ.get("DLX_BDROM") or f"/media/{getpass.getuser()}/BDROM"
STREAM_DIR = f"{BDROM}/BDMV/STREAM"
# 8 MHz / {512, 768, 1024}. The chip has no other rates and 15,625 is the one
# every budget in this project is written against (FINDINGS 52).
RATES = {15625: 512, 10417: 768, 7813: 1024}
def extract(stream, out, rate=15625, start=None, dur=None):
if rate not in RATES:
raise SystemExit(f"{rate} is not an MSM6258 rate: {sorted(RATES)}")
src = f"{STREAM_DIR}/{stream}.m2ts"
cmd = ["ffmpeg", "-v", "error"]
if start is not None: cmd += ["-ss", str(start)]
if dur is not None: cmd += ["-t", str(dur)]
cmd += ["-i", src, "-vn", "-ac", "1", "-ar", str(rate),
"-f", "s16le", "-acodec", "pcm_s16le", out, "-y"]
subprocess.check_call(cmd)
n = os.path.getsize(out) // 2
print(f"{stream}: {n} samples @ {rate} Hz mono = {n/rate:.3f} s -> {out}")
return n
if __name__ == "__main__":
# extract_audio.py <stream> <out.raw> [rate] [start_s] [dur_s]
stream, out = sys.argv[1], sys.argv[2]
rate = int(sys.argv[3]) if len(sys.argv) > 3 else 15625
start = float(sys.argv[4]) if len(sys.argv) > 4 else None
dur = float(sys.argv[5]) if len(sys.argv) > 5 else None
extract(stream, out, rate, start, dur)