Claire's writing — 8 documents, 33 snapshots, 103 suggestions, 3 vocabulary
words and an image — now belongs to her account rather than to the pre-auth
'local' user, and the VPS is canonical. millenia was left running and
untouched as a frozen fallback; it diverges the moment either side is
written to, so it wants retiring rather than syncing.
The plan's stated prerequisite, that she log in once so her subject exists,
turned out to be false. Authentik's hashed_user_id sub is the user's uid,
derived from her id and the instance secret, so it can be read in advance —
which means the data moves first and she signs in to find her writing
already there, instead of to an empty Petal that fills in later.
The fix here is to the liveness guard, and it is the second attempt at it.
PRAGMA locking_mode = EXCLUSIVE goes on holding its lock after being set
back to NORMAL — SQLite only lets go on that connection's next database
access — so against a real WAL database the script locked itself out of its
own VACUUM INTO backup. It passed locally because the test database had
come out of VACUUM INTO and so was never in WAL mode: the fixture didn't
look like production, the same way the stub identity provider's slashless
issuer didn't. The probe now runs on its own connection and closes it.
Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
The gate existed because Petal authenticated nobody and a public hostname
was therefore a public, writable API. It no longer is: every /api route
answers 401 without a session, so the only thing an anonymous visitor
reaches is the app shell and its redirect to Authentik. The separate
unauthenticated /api/health router goes with it — it only existed to escape
the middleware. A second password in front of a real login is one more
thing to lose.
Also records the two problems this deployment actually hit, since both fail
before the login page appears and neither is obvious from the error: the
issuer's trailing slash is significant, and a provider created through the
API rather than the admin UI comes up with an empty grant_types.
Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
Petal is now an OIDC client in its own right rather than trusting a header
from the proxy. The Phase-0 Resolver seam was the only integration point:
main.go picks the session store when Authentik is configured and the static
local user otherwise, and no handler or query moved for either.
internal/auth gains three pieces. session.go issues an opaque cookie token
and stores only its SHA-256, so a database copy yields nothing usable; the
30-day expiry slides on every request, throttled to one write an hour, and
logout deletes the row rather than just the cookie. oidc.go runs the
authorization-code flow with state, nonce and PKCE, and discovers the
provider lazily and on retry — an Authentik outage should block new logins
without stopping Petal booting or invalidating live sessions. users.go
provisions accounts from the token's claims and gates them on an allowlist
that matches emails as well as subject ids, since a subject is an opaque
uuid that doesn't exist until someone has already logged in once.
Migration 0010 lands sessions, images and users.pair_lang together. The
images table closes the capability-URL hole the Phase-0 audit flagged: a
hash was previously enough to fetch anyone's picture. Rows are keyed
(name, user_id) so one file can have several owners and deduplication
survives; a stranger gets 404 rather than 403, the cache header drops to
private, and files already on disk are claimed at startup or every image
already pasted into a document would 404.
On the frontend a single 401 interceptor feeds a warm bilingual sign-in
overlay, drawn over a still-visible editor because nothing has been taken
away. Behind it is the part that matters: a save that comes back 401
stashes its body to localStorage before anything else and stops the
auto-save loop, and reopening that document after signing in merges the
draft back and saves it. An expired session must not cost writing.
Writing the round-trip test against a stub identity provider turned up a
real bug: the one-shot state/nonce/PKCE cookies were cleared in a defer,
which runs after the redirect has written the response header, so the
clearing Set-Cookie was silently dropped and they lingered for their full
ten minutes.
Also swaps the emoji favicon for a drawn sakura, which renders as Petal's
own rose palette everywhere instead of whatever each platform's font
decides, and doubles as the app tile in Authentik.
Migration 0010 verified against a VACUUM INTO copy of the live millenia
database: counts intact, FTS still matching, the one existing image
claimed.
Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
The canonical instance -- the one with her actual writing -- turned out
to be the least protected thing in the estate:
- No scheduled backup at all; the newest snapshot was a month old. Now
petal-backup.timer: VACUUM INTO, gzip, age-encrypt with the parodia
public recipient, push to the VPS over headscale with a size check,
prune both ends. Persistent=true because the box is not on 24/7.
Neither machine can decrypt what it holds; the identity is offline.
- Petal ran as a bare ./petal with PPID 1, so a crash or reboot left it
down until somebody noticed. Now petal.service, verified by kill -9.
- The Piper units retried forever without ever failing: RestartSec=3
against systemd's default 10s window means the burst limit is never
reached, which is how a dead service logged 26,800+ restarts over a
day while read-aloud silently fell back to browser speech.
StartLimitIntervalSec=300 makes a broken Piper show up in --failed.
backup-petal.sh now handles both deployment shapes (compose exec on the
VPS, local binary on millenia) and encrypts before anything leaves the
host. The VPS no longer uses it -- Petal rides parodia-backup there.
Petal's data directory (petal.db, uploaded images, and the TTS cache --
which is synthesized audio of her sentences) now lives on a LUKS volume.
LUKS-on-a-file rather than gocryptfs because Petal is SQLite in WAL mode:
WAL needs a shared-memory index mapped consistently across processes, and
FUSE has a long history of subtle mmap/locking differences. A block
device with ext4 behaves exactly like a disk to SQLite.
The key sits on the same host, which is a deliberate availability
tradeoff and is documented as such: it stops a decommissioned disk or a
raw block-device read, not anyone holding the whole VM image.
Two bugs found by rehearsing a reboot rather than trusting the setup:
- Mounting over a directory HIDES its contents, it does not remove them.
The first run left the original plaintext petal.db and WAL sitting on
the unencrypted root filesystem, invisible under the mount -- exactly
what the exercise was meant to eliminate. Now shredded before the
mount, with a refusal if the mountpoint will not come up empty.
- systemd-cryptsetup was not installed on the host, so /etc/crypttab was
being ignored entirely and the volume would never have unlocked at
boot. The script now refuses to run without the generator present.
Rebinding vllm-chat was the expensive option: Petal, Gogobee and Open
WebUI all point at 127.0.0.1:8000, and Open WebUI keeps its endpoint in
its own database rather than in env, so moving the bind address meant
editing three consumers and reloading a 35B AWQ model. The socat unit
adds a second listener on the headscale address instead -- local callers
untouched, no downtime, and the only new exposure is on the VPN. Bound
to 100.64.0.2 specifically, never 0.0.0.0: the far end is a public host.
Verified: a grammar checkpoint from petal.parodia.dev returns real
suggestions in ~3s over the VPN.
Also documents two things found on millenia that were invisible from
outside it:
- Piper had been dead since the Jul 26 reboot, 26,800+ failed restarts,
with read-aloud silently falling back to browser Web Speech. An OS
upgrade moved /usr/bin/python3 from 3.13 to 3.14, and venv/bin/python3
is a symlink to the system interpreter, so lib/python3.13/site-packages
went invisible -- sys.path had no site-packages at all. Recreating the
venv lands Piper 1.6.0, which is what TTS_PATH exists for.
- The canonical instance has no automated backup (newest snapshot a
month old) and runs unsupervised with PPID 1. Both written down; the
backup one is pending the encryption-at-rest decision.
Backups on the VPS now ride parodia-backup (age-encrypted, offsite, S3),
using VACUUM INTO rather than that script's iterdump helper -- iterdump
does not reproduce an FTS5 virtual table, so a restore would have come
back with cross-document search silently missing.
deploy/README.md becomes the real runbook: the VPS stack, Traefik, the
headscale LLM link, the interim edge gate, backups and restore. The
millenia Piper notes move to an appendix -- that instance still runs
them, and it is still canonical.
Two items are called out as outstanding rather than done, because both
need access to millenia: vLLM is not bound to its headscale interface,
so no AI pass works from the VPS yet, and parodia's ssh key is not
authorized there, so backups are VPS-local only -- which is not a backup
in the sense that matters. Each has its one-command fix written down.
Also folds in DreamDict gaining Spanish: es was explicitly gated on that
dataset existing, so it moves from "Later / not now" to a normal
follow-on pair after pt-PT, and the Phase 20 provider seam should cover
it from the start.
Petal authenticates nobody yet -- StaticResolver hands every request the
same local user -- so on a public host the whole API is open: anyone who
finds the hostname can read and write documents and fill the disk with
image uploads. Traefik holds the door until the OIDC flow exists.
/api/health keeps its own higher-priority router with no middleware, so
the acceptance criterion (public health endpoint, reachable by the
monitoring on this box) still holds. Both the middleware and that router
are deleted when Phase 16 lands.
piper-tts 1.6.0 moved synthesis from POST / to POST /synthesize, with an
identical request body; the VPS sidecars run 1.6.0 and returned 405 to
every read-aloud request, while millenia's older server still expects /.
Rather than pinning both deployments to one Piper release, the path is
configuration -- default "/" keeps millenia working untouched, and the
compose stack sets /synthesize. The container healthcheck moves with it,
since it was probing the old route too.
The image's own petal user (uid 10001) has no claim on a bind-mounted
host directory, so SQLite came up with "unable to open database file
(14)" and the container restart-looped. Run as the stack directory's
owner instead of chowning ./data to 10001 -- the backup script gzips
snapshots in place from the host, so that account needs write access to
the same directory. Still non-root.
Deploy plumbing so Petal can run on the public VPS behind the Traefik
already on that box, with vLLM reached over headscale.
- Dockerfile: node build -> go build -> alpine runtime. CGO stays off
(modernc SQLite is pure Go), so the runtime layer exists only for
ffmpeg (read-aloud transcodes Piper's WAV) and tzdata (the companion's
bedtime nag and night mode read the local clock). Runs as uid 10001
with /data as the single writable mount.
- docker-compose.yml: Traefik labels following this host's convention
(external `traefik` network, `web-secure` entrypoint, `default` cert
resolver). Petal publishes no host port. ./data is a bind mount, not a
named volume, so the nightly backup and a restore are reachable from
the host.
- Piper runs as two sibling containers rather than host systemd units.
The plan assumed Piper was already installed on the VPS; it is not,
the host has no lingering user session to keep user units alive, and
containers keep the TTS ports on an internal network unreachable from
anywhere but Petal. One image, voice chosen per service, model cached
in a shared volume -- so the pt-PT voice is a new service, not a new
image.
- db.Backup + a `-backup` flag: VACUUM INTO, not a file copy. Petal runs
in WAL mode, so the newest committed pages may live in petal.db-wal;
copying the three files separately can capture a torn mid-checkpoint
state. VACUUM INTO reads one coherent snapshot without taking a write
lock, and emits a single file with no -wal/-shm companions. Refuses an
existing destination so a failed run can't destroy the last good
backup.
- deploy/backup-petal.sh: nightly snapshot, compress, push to millenia
over headscale with a post-transfer size check, prune both sides.
- deploy/petal.env.example: LLM_TIMEOUT raised 30s -> 90s for the
WAN+VPN round trip, since the voice and collocation passes send a
whole document and the timeout is a hard deadline on Complete.
Replace the browser's robotic Web Speech API (espeak on Chromium) as the
primary read-aloud path with server-side Piper neural TTS, served by petal
and kept fully offline on millenia next to Ollama.
- internal/tts: proxy short passages to Piper, transcode WAV -> mp3/opus via
ffmpeg, content-addressed disk cache (instant re-taps). Each Piper instance
loads one voice, so language routes to its own endpoint:{voice}. Unknown
language -> 404 so the client falls back to Web Speech. UTF-8-safe truncation
for multibyte (Chinese) text.
- config: TTS_ENDPOINT / TTS_ENDPOINT_ZH / TTS_VOICE_EN / TTS_VOICE_ZH /
TTS_CACHE_DIR / TTS_TIMEOUT / TTS_AUDIO_FORMAT. Route mounts only when
TTS_ENDPOINT is set; otherwise unchanged behavior.
- web/audio/speech.ts: speak() hits /api/tts first, falls back to Web Speech on
any failure; rapid-tap-safe via a request token. Call sites unchanged.
- deploy/: Piper user systemd units (EN :5005, ZH :5006), setup script, README.
English (en_US-amy-medium) and Chinese (zh_CN-huayan-medium) are both live.
Claude-Session: https://claude.ai/code/session_016Yr6jELuRc7hyzYLccQKZd