Find the sustained action sequence: it breaks both profiles
The open risk since session 2 was "a sustained action sequence could still break the bitrate", with every clip measured so far being 1.2-1.7 s. Closed by measurement rather than by sampling clips by hand. 07_motion_survey.py scans a whole stream at 96x72 for the hottest sliding window of inter-frame difference. On 00223 the spread between the quietest and hottest sustained 10 s windows is 10.6x, which is the argument for not eyeballing it. Hottest is t=539.4s, the Singe endgame. There, with the fixed lam the CLI uses, sasi overshoots 110 -> 129.6 KB/s (+18%) and scsi 280 -> 373.8 KB/s (+34%). Rate control moves from "insurance, not a fix" to required, and is promoted above the full-disc survey. The bus is not broken -- 381.6 KB/s still fits the 488 KB/s figure -- so FINDINGS 21 survives, at 78% of the pipe instead of a comfortable margin. Three further corrections fall out: - The two largest streams on the disc are bonus material. 00216 is the feature with a burned-in commentary PiP; 00215 is the commentary. 00223 is the clean 9.4 min. A size-ranked survey would have encoded live action. - On hard content the 256-colour scene palette (31.33 dB) binds well before the X68000 display (40.81 dB); scsi is already within 0.51 dB of it. - FINDINGS 24.5's architecture question resolves to "both paths, chosen per frame": 30-53% of frames sit above the 70% crossover. Picking per frame costs a median 37.0% of the frame budget and caps at 53.6%. Reporting for this is wired into encode.py, which previously only printed a mean over all frames -- the one statistic that cannot answer a per-frame question. extract.py takes optional start/dur; 08_mode_map.py renders source | decoded | block-mode map to .webm. Claude-Session: https://claude.ai/code/session_01194oWYW8DQXK1SZ2DnChW6
This commit is contained in:
@@ -0,0 +1,57 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Find the worst sustained-motion window in a stream, cheaply.
|
||||
|
||||
STATUS lists "a sustained action sequence is the one thing that could still
|
||||
break the bitrate" as the open risk, and every clip measured so far has been
|
||||
1.2-1.7 s. Picking a hot clip by eye is how you get a comfortable answer, so
|
||||
this scans the whole stream instead.
|
||||
|
||||
Proxy: mean absolute inter-frame difference at 96x72, decimated to the target
|
||||
12 fps. It is a proxy, not a bitrate -- but the codec's cost is dominated by
|
||||
how many blocks fail SKIP, and that is what frame difference measures. The
|
||||
window it picks then gets encoded for real.
|
||||
|
||||
Usage: python3 tools/analysis/07_motion_survey.py 00223 [window_seconds]
|
||||
"""
|
||||
import subprocess, sys
|
||||
import numpy as np
|
||||
|
||||
STREAM_DIR = "/media/reala-misaki/BDROM/BDMV/STREAM"
|
||||
W, H, FPS = 96, 72, 12
|
||||
|
||||
def frames(stream):
|
||||
src = f"{STREAM_DIR}/{stream}.m2ts"
|
||||
vf = f"fps={FPS},crop=1440:1080:240:0,scale={W}:{H}:flags=bilinear"
|
||||
p = subprocess.Popen(["ffmpeg", "-v", "error", "-i", src, "-vf", vf,
|
||||
"-f", "rawvideo", "-pix_fmt", "gray", "-"],
|
||||
stdout=subprocess.PIPE)
|
||||
buf = p.stdout.read()
|
||||
p.wait()
|
||||
n = len(buf) // (W * H)
|
||||
return np.frombuffer(buf[:n*W*H], np.uint8).reshape(n, H, W).astype(np.int16)
|
||||
|
||||
def main():
|
||||
stream = sys.argv[1]
|
||||
win_s = float(sys.argv[2]) if len(sys.argv) > 2 else 10.0
|
||||
f = frames(stream)
|
||||
d = np.abs(np.diff(f, axis=0)).mean(axis=(1, 2)) # per-frame motion energy
|
||||
print(f"{stream}: {len(f)} frames @ {FPS}fps = {len(f)/FPS:.1f}s")
|
||||
print(f" motion energy mean {d.mean():.2f} median {np.median(d):.2f} "
|
||||
f"p90 {np.percentile(d,90):.2f} max {d.max():.2f}")
|
||||
|
||||
w = int(win_s * FPS)
|
||||
if len(d) < w:
|
||||
print("stream shorter than the window"); return
|
||||
# sustained = highest mean over a sliding window, not the single hottest frame
|
||||
k = np.convolve(d, np.ones(w) / w, mode="valid")
|
||||
best = int(np.argmax(k))
|
||||
print(f" hottest sustained {win_s:.0f}s window: t = {best/FPS:.1f}s "
|
||||
f"(mean {k[best]:.2f}, {k[best]/d.mean():.2f}x stream mean)")
|
||||
quiet = int(np.argmin(k))
|
||||
print(f" quietest {win_s:.0f}s window: t = {quiet/FPS:.1f}s "
|
||||
f"(mean {k[quiet]:.2f}, {k[quiet]/d.mean():.2f}x stream mean)")
|
||||
np.save(f"tmp/motion_{stream}.npy", d)
|
||||
print(f" per-frame energy -> tmp/motion_{stream}.npy")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,107 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Render what the codec is actually DOING, per block, per frame.
|
||||
|
||||
Three panels at 12 fps: the palettised source (the real quality ceiling, not
|
||||
1080p -- FINDINGS 11), the decoded output, and a block-mode map.
|
||||
|
||||
The mode map is not decoration. FINDINGS 24.5 makes the per-frame non-SKIP
|
||||
fraction the number that selects the decoder's inner loop, and a percentile
|
||||
cannot show you that the non-SKIP blocks are CLUSTERED (a moving character on a
|
||||
held background) rather than scattered. Clustering is what a run-length over
|
||||
the mode headers would exploit.
|
||||
|
||||
SKIP left as the previous frame, costs the 68000 nothing
|
||||
V1 one 4x4 codeword, 1 byte
|
||||
V4 four 2x2 codewords, 4 bytes
|
||||
RAW 16 literal palette indices -- the escape that makes lam=0 pixel-exact
|
||||
|
||||
Usage: python3 tools/analysis/08_mode_map.py <frames_dir> <out.webm> [--profile p]
|
||||
[--scale N]
|
||||
Output format follows the extension. Prefer .webm: GIF re-quantises to 256
|
||||
colours, which is a poor fit for output whose subject is colour fidelity.
|
||||
"""
|
||||
import sys, os
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "encoder"))
|
||||
import numpy as np
|
||||
from PIL import Image
|
||||
import vq as VQ, vq_hybrid as H, ratectl as RC
|
||||
|
||||
MODE_RGB = np.array([[ 20, 22, 30], # SKIP - near black, costs nothing
|
||||
[ 60, 150, 230], # V1 - blue
|
||||
[ 80, 200, 120], # V4 - green
|
||||
[235, 90, 70]], # RAW - red, the expensive escape
|
||||
dtype=np.uint8)
|
||||
LABEL = ["SKIP", "V1", "V4", "RAW"]
|
||||
|
||||
SCALE = 1
|
||||
LOSSLESS = False
|
||||
|
||||
def main():
|
||||
global SCALE, LOSSLESS
|
||||
LOSSLESS = "--lossless" in sys.argv
|
||||
src, out = sys.argv[1], sys.argv[2]
|
||||
if "--scale" in sys.argv:
|
||||
SCALE = int(sys.argv[sys.argv.index("--scale")+1])
|
||||
prof = RC.PROFILES[sys.argv[sys.argv.index("--profile")+1]
|
||||
if "--profile" in sys.argv else "sasi"]
|
||||
m = H.build(src, k1=prof["k1"], k4=prof["k4"])
|
||||
enc = H.encode(m, lam=prof["lam"])
|
||||
pal, H_, W_ = m["pal"], m["H"], m["W"]
|
||||
nbx, nby = W_ // 4, H_ // 4
|
||||
|
||||
frames, stats = [], []
|
||||
for f, (rec, mode) in enumerate(zip(enc["recon"], enc["modes"])):
|
||||
srcp = pal[m["idx"][f]]
|
||||
decp = pal[rec]
|
||||
mmap = MODE_RGB[mode.reshape(nby, nbx)].repeat(4, 0).repeat(4, 1)
|
||||
# tint the mode map with the decoded luma so the action stays legible
|
||||
luma = decp.mean(2, keepdims=True) / 255.0
|
||||
mmap = (mmap * (0.45 + 0.55 * luma)).astype(np.uint8)
|
||||
gap = np.full((H_, 3, 3), 60, np.uint8)
|
||||
panel = np.hstack([srcp, gap, decp, gap, mmap])
|
||||
im = Image.fromarray(panel)
|
||||
if SCALE != 1:
|
||||
im = im.resize((panel.shape[1]*SCALE, H_*SCALE), Image.NEAREST)
|
||||
frames.append(im)
|
||||
stats.append([(mode == i).mean() for i in range(4)])
|
||||
|
||||
if out.endswith(".webm") or out.endswith(".mp4"):
|
||||
# Preferred. GIF would impose its own 256-colour palette on top of
|
||||
# output whose entire subject is colour fidelity, and costs ~4x the
|
||||
# bytes doing it. -lossless keeps the panels pixel-exact.
|
||||
import subprocess
|
||||
w, h = frames[0].size
|
||||
cmd = ["ffmpeg", "-v", "error", "-y", "-f", "rawvideo", "-pix_fmt", "rgb24",
|
||||
"-s", f"{w}x{h}", "-r", "12", "-i", "-"]
|
||||
# yuv444p, not 420: the mode map is flat saturated colour on a 4-pixel
|
||||
# grid, and chroma subsampling smears exactly those edges. Lossless is
|
||||
# available but runs larger than the GIF on this content; crf 18 in 444
|
||||
# is visually clean at a quarter the size.
|
||||
if out.endswith(".webm"):
|
||||
cmd += ["-c:v", "libvpx-vp9", "-pix_fmt", "yuv444p", "-row-mt", "1"]
|
||||
cmd += ["-lossless", "1"] if LOSSLESS else ["-crf", "18", "-b:v", "0"]
|
||||
else:
|
||||
cmd += ["-c:v", "libx264", "-crf", "12", "-pix_fmt", "yuv444p"]
|
||||
p = subprocess.Popen(cmd + [out], stdin=subprocess.PIPE)
|
||||
for f in frames:
|
||||
p.stdin.write(np.asarray(f.convert("RGB")).tobytes())
|
||||
p.stdin.close(); p.wait()
|
||||
else:
|
||||
# One shared adaptive palette: per-frame palettes are what make a naive
|
||||
# GIF of this enormous, and a stable palette also stops the mode-map
|
||||
# colours shimmering between frames.
|
||||
shared = frames[0].quantize(colors=192, method=Image.MEDIANCUT)
|
||||
q = [f.quantize(palette=shared, dither=Image.NONE) for f in frames]
|
||||
q[0].save(out, save_all=True, append_images=q[1:],
|
||||
duration=1000//12, loop=0, optimize=True)
|
||||
st = np.array(stats)
|
||||
print(f"{len(frames)} frames -> {out} ({os.path.getsize(out)/1024:.0f} KB)")
|
||||
print(" panels: palettised source | decoded | block mode map")
|
||||
for i, n in enumerate(LABEL):
|
||||
print(f" {n:4s} mean {100*st[:,i].mean():5.1f}% "
|
||||
f"per-frame range {100*st[:,i].min():5.1f}% .. {100*st[:,i].max():5.1f}%")
|
||||
ns = 100 * (1 - st[:, 0])
|
||||
print(f" non-SKIP: median {np.median(ns):.1f}% p90 {np.percentile(ns,90):.1f}%")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -31,6 +31,12 @@ sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
import numpy as np
|
||||
import vq as VQ, vq_hybrid as H, ratectl as RC
|
||||
|
||||
# Measured on the emulated 68000, FINDINGS 24. Instruction cycles against
|
||||
# zero-wait-state memory, so these are floors, not hardware predictions.
|
||||
BLIT_PCT = 53.6 # V1: compose in RAM, then a row-linear movem.l blit
|
||||
DIRECT_PCT = 76.6 # V4: write every block straight into GVRAM
|
||||
CROSSOVER_PCT = 100 * BLIT_PCT / DIRECT_PCT
|
||||
|
||||
|
||||
def pack_modes(mode):
|
||||
"""2 bits per block, MSB-first -- cheap for the 68000 to shift out."""
|
||||
@@ -133,6 +139,24 @@ def main():
|
||||
print(f" modes: SKIP {r['skip']:.1f}% V1 {r['v1']:.1f}% "
|
||||
f"V4 {r['v4']:.1f}% RAW {r['raw']:.1f}%")
|
||||
|
||||
# PER-FRAME non-SKIP distribution. The mean above cannot answer the
|
||||
# decoder-architecture question (FINDINGS 24.5): decode-direct-to-GVRAM
|
||||
# costs 76.6% of a 12fps frame budget x (non-SKIP fraction), while
|
||||
# compose-in-RAM-then-blit is a flat 53.6% regardless. They cross at 70%,
|
||||
# and that is a decision taken FRAME BY FRAME -- a scene cut is ~100%
|
||||
# non-SKIP and a held frame near 0%, so their mean describes no real frame.
|
||||
ns = np.array([100 * (mm != 0).mean() for mm in enc["modes"]])
|
||||
over = int((ns > CROSSOVER_PCT).sum())
|
||||
print(f" non-SKIP blocks/frame: median {np.median(ns):.1f}% "
|
||||
f"p90 {np.percentile(ns, 90):.1f}% max {ns.max():.1f}%")
|
||||
print(f" frames above the {CROSSOVER_PCT:.0f}% blit crossover: "
|
||||
f"{over}/{len(ns)} ({100*over/len(ns):.1f}%) -> "
|
||||
f"{'compose+blit wins on those' if over else 'direct-to-GVRAM wins throughout'}")
|
||||
cost = np.minimum(BLIT_PCT, DIRECT_PCT * ns / 100)
|
||||
print(f" display cost if the player picks the cheaper path per frame: "
|
||||
f"median {np.median(cost):.1f}% p90 {np.percentile(cost, 90):.1f}% "
|
||||
f"max {cost.max():.1f}% of a 12fps frame")
|
||||
|
||||
if a.preview:
|
||||
from PIL import Image
|
||||
f = len(idx) // 2
|
||||
|
||||
@@ -35,7 +35,13 @@ def extract(stream, outdir, fps=12, mode="crop", start=None, dur=None):
|
||||
return n
|
||||
|
||||
if __name__ == "__main__":
|
||||
# extract.py <stream> <outdir> [fps] [mode] [start_s] [dur_s]
|
||||
# start/dur matter for the long streams: 00223 is 9.4 min and the parts
|
||||
# that stress the codec are a few seconds each (tools/analysis/07 finds
|
||||
# them). Without them a survey silently averages action into idle scenes.
|
||||
stream, outdir = sys.argv[1], sys.argv[2]
|
||||
fps = int(sys.argv[3]) if len(sys.argv) > 3 else 12
|
||||
mode = sys.argv[4] if len(sys.argv) > 4 else "crop"
|
||||
extract(stream, outdir, fps, mode)
|
||||
start = float(sys.argv[5]) if len(sys.argv) > 5 else None
|
||||
dur = float(sys.argv[6]) if len(sys.argv) > 6 else None
|
||||
extract(stream, outdir, fps, mode, start, dur)
|
||||
|
||||
Reference in New Issue
Block a user