Phase 24, the fr half: langpack, Hunspell dictionary, Piper voice, and the
lexicon coverage that turned out to have been measured already (63.1%, better
than pt-PT's 62.1%). No migration; not deployed.
The plan recorded that build_ptpt_dictionary.py "generalizes" to French. It
did not. It handled single-character flags and plain PFX/SFX and stopped on
everything else, and fr.aff uses four of the things it stopped on. FLAG long
is the dangerous one: French flags are two characters, so the old reader's
set(flagstr) yields a bag of unrelated letters and expands every entry through
the wrong paradigm without ever erroring. Plus continuation flags (French
really does affix an affixed form), NEEDAFFIX on 68,075 of 84,140 stems, and
FULLSTRIP. Renamed build_hunspell_dictionary.py with a per-language profile,
asserting that CIRCUMFIX and FORBIDDENWORD are still unused rather than
assuming it — and it rebuilds pt-PT byte-identical to the shipped asset, which
is the only thing that makes "generalized" a claim rather than a hope.
Elision was decided by building both halves and measuring. Keeping l'arbre and
its thirty-three siblings: 3,159,832 forms, 8.25 MB gzipped. Dropping them:
473,326 and 1.19 MB. They are not new words, but the tokenizer keeps internal
apostrophes, so they genuinely would have been underlined — so they moved out
of the dictionary into withElision, which splits at a known clitic and still
requires the remainder to be a word (l'zzzz stays flagged). Real nspell: 369 ms
and 74 MB, against pt-PT's 842 ms and 139 MB, on the larger language.
Where the regional trap lives is the mirror image of Portuguese's: every fr_*
Piper voice is fr_FR and Debian's fr_FR/fr_CA/fr_BE dictionaries are one shared
word list, so nothing can be quietly wrong about the country and the whole
decision sits in the copy. What French has instead is the 1990 reform, packaged
three ways; comprehensive ships, because Petal never corrects her French and
coût and cout are both correct.
Then the interim review pass, at the user's suggestion and explicitly "for
now": four models read each Latin pack independently, and only findings at
least two of them reached on their own were applied — five per pack. It earned
its keep on the pack that was already live. pt-PT was carrying pre-Acordo
spellings (adjectivos, actualmente) in a file whose own header commits to
post-Acordo, plus Brazilian decepção, because the Phase 21 greps checked for
Brazilian vocabulary and never checked the pack against its own spelling
policy. That grep now exists and was confirmed to fail on the old text before
being kept. Where reviewers agreed a line was wrong but split on the fix, the
wording is mine and the reasoning is in BUILD_PLAN rather than averaged away.
Still owed, and both packs now say so precisely: a quorum of models agreeing is
agreement, not authority. No native speaker has read either pack, and none of
this has been seen in a browser.
go build/vet/test clean, tsc, vite build, vitest 190/190.
Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
Phase 21's infra half. Two things the pt-PT pair needs from TTS, and one
thing every learner has wanted since Phase 11.
**A language is no longer a code change.** The handler knew exactly two
languages, named in the Config struct: English on TTS_ENDPOINT and Chinese
on TTS_ENDPOINT_ZH. Petal now discovers its Piper instances from the
environment — English keeps the unsuffixed pair it has always had, and
every other language is a TTS_ENDPOINT_<LANG>/TTS_VOICE_<LANG> pair — so
fr and es cost a compose service and two lines of .env. <LANG> is the base
tag, because an environment variable name cannot hold pt-PT's hyphen and
only one Portuguese model is loaded either way. A language configured by
halves is dropped rather than routed: half a configuration should reach
the client as "no voice here, use Web Speech", not as an instance that
errors on every tap. The startup line now names the voices it actually
resolved rather than the English endpoint it was handed — the same lesson
the dictionary line learned last week.
**pt_PT-tugão-medium is the only European voice Piper ships.** The other
five pt models in the catalogue are Brazilian, so the default anyone
reaches for is the wrong country — the same trap as `dictionary-pt`
packaging VERO, arriving through the catalogue rather than through the
model. Named explicitly in compose, with the query that checks it in the
deploy README.
**The slow replay** (SUGGESTIONS §5e) is `slow: true` on /api/tts, raising
Piper's length_scale to ~4/3. Piper stretches durations rather than
resampling, so it stays a voice instead of a groan. The pace is part of
the cache key — without it the slow replay of a word already heard at
normal speed would be served back at normal speed, which is the one
request where the difference is the whole point. 🐢 sits beside 🔊 on the
word card, the selection bubble and the garden flashcard; the Web Speech
fallback slows too, so the button means the same thing when Piper is down.
**And the other reading gets her own voice.** The `alsoIn` block — the
Portuguese sense of a word that is also English — now speaks in the pair's
locale, which the pack names (`locale`) rather than anything inferring it
from the letters. "comum" is spelled identically in both halves; a
detector would have to guess, and this is the same reason the gloss shows
both directions instead of picking one.
Tests: config discovery (both existing deployment shapes, half-configured
languages dropped, the pre-map voice defaults preserved), the slow scale
and its separate cache entry, pt routing on the base tag with pt-BR
landing on the European instance, and speech.ts's request body. The i18n
shape suite now asserts every pack names a speakable locale in its own
language — and that pt-PT's is not pt-BR.
Verified: go build/vet/test, tsc, vitest 125/125, vite build. Live smoke
against two fake Piper servers: en/pt × normal/slow all reached the right
instance at the right length_scale with four distinct cache entries, and
an unconfigured language still 404s.
Word lookups now come from DreamDict's dict.db for every pair but Chinese —
opened read-only beside petal.db, no service, nothing over the VPN, because a
hover gloss has to answer in milliseconds.
`Provider` is the two questions the popover and the tooltip already asked, so
the embedded *Lexicon satisfies it with no changes at all; Set.For(lang) is the
single place the choice between them is made. The prerequisite in the dreamdict
repo turned out to be two things, not one: the module path was unfetchable
*and* the query layer sat in internal/, which no other module may import
whatever the module is called. Both fixed upstream.
The plan's central assumption did not survive the data. It mapped
Gloss ← Translate(word, "en", L1) one-to-one; against the real 452 MB database
that table answers for 17% of the 2,000 commonest English words into pt-PT.
Wiktionary's translation sections are thin in that direction — "ephemeral",
"think" and "quickly" have no en→pt-PT row at all. Shared WordNet synsets
answer for 61%, so DreamDict gained Equivalents() and Petal glosses through it.
Ordering those was wrong in an instructive way too: sorting by frequency
glosses "think" as lembrar, "remember", because lembrar is the commoner
Portuguese word even though pensar shares six of think's synsets to lembrar's
one. Counting sense agreement first asks the right question.
The same measurement is why zh stays on ECDICT: DreamDict reaches a Chinese
gloss for 53% of those words, ECDICT for nearly all of them. The plan said
converge only if quality holds. It didn't, so nothing converged.
Two decisions about failure worth keeping. A missing dict.db is not an error —
a laptop checkout has never had one — but a present-and-never-imported one is,
because that is a half-finished deploy. And a pt-PT writer with no dictionary
falls back to the embedded datasets with the gloss suppressed, keeping
definitions, synonyms and phonetics rather than blanking the popover: an empty
field reads as "not found", the wrong language reads as broken.
The new fields surface as an etymology line and a three-band chip. Three, not
five: the difficulty score separates "everyday" from "you'll have to explain
this" but cannot rank obfuscate against serendipity, and a finer scale would be
a confident-looking lie. An unscored word gets no chip.
Writing the tests found two bugs first — trimEtymology sliced by byte, which
would have emitted invalid UTF-8 for exactly the Greek and Latin etymologies
the feature exists for, and its ellipsis path overran its own cap.
go build/vet/test, tsc, vite, vitest 96/96 clean; live smoke against the real
dict.db with one instance flipped from zh to pt-PT mid-run.
Not deployed: go.mod still replaces github.com/prosolis/dreamdict with
../dreamdict, so the Docker build needs the two upstream commits pushed and the
replace dropped. The deployed dict.db also predates DreamDict's Spanish data.
Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
The gate existed because Petal authenticated nobody and a public hostname
was therefore a public, writable API. It no longer is: every /api route
answers 401 without a session, so the only thing an anonymous visitor
reaches is the app shell and its redirect to Authentik. The separate
unauthenticated /api/health router goes with it — it only existed to escape
the middleware. A second password in front of a real login is one more
thing to lose.
Also records the two problems this deployment actually hit, since both fail
before the login page appears and neither is obvious from the error: the
issuer's trailing slash is significant, and a provider created through the
API rather than the admin UI comes up with an empty grant_types.
Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
The mountpoint directory exists whether or not the encrypted volume is
mounted, so a boot where the unlock failed would start Petal against an
empty unencrypted directory and serve a blank database -- the failure
mode that looks like data loss. .volume-ok lives on the encrypted
filesystem and is bind-mounted with create_host_path:false, so its
absence is a container start failure instead of a silent empty DB.
Petal authenticates nobody yet -- StaticResolver hands every request the
same local user -- so on a public host the whole API is open: anyone who
finds the hostname can read and write documents and fill the disk with
image uploads. Traefik holds the door until the OIDC flow exists.
/api/health keeps its own higher-priority router with no middleware, so
the acceptance criterion (public health endpoint, reachable by the
monitoring on this box) still holds. Both the middleware and that router
are deleted when Phase 16 lands.
piper-tts 1.6.0 moved synthesis from POST / to POST /synthesize, with an
identical request body; the VPS sidecars run 1.6.0 and returned 405 to
every read-aloud request, while millenia's older server still expects /.
Rather than pinning both deployments to one Piper release, the path is
configuration -- default "/" keeps millenia working untouched, and the
compose stack sets /synthesize. The container healthcheck moves with it,
since it was probing the old route too.
The image's own petal user (uid 10001) has no claim on a bind-mounted
host directory, so SQLite came up with "unable to open database file
(14)" and the container restart-looped. Run as the stack directory's
owner instead of chowning ./data to 10001 -- the backup script gzips
snapshots in place from the host, so that account needs write access to
the same directory. Still non-root.
Deploy plumbing so Petal can run on the public VPS behind the Traefik
already on that box, with vLLM reached over headscale.
- Dockerfile: node build -> go build -> alpine runtime. CGO stays off
(modernc SQLite is pure Go), so the runtime layer exists only for
ffmpeg (read-aloud transcodes Piper's WAV) and tzdata (the companion's
bedtime nag and night mode read the local clock). Runs as uid 10001
with /data as the single writable mount.
- docker-compose.yml: Traefik labels following this host's convention
(external `traefik` network, `web-secure` entrypoint, `default` cert
resolver). Petal publishes no host port. ./data is a bind mount, not a
named volume, so the nightly backup and a restore are reachable from
the host.
- Piper runs as two sibling containers rather than host systemd units.
The plan assumed Piper was already installed on the VPS; it is not,
the host has no lingering user session to keep user units alive, and
containers keep the TTS ports on an internal network unreachable from
anywhere but Petal. One image, voice chosen per service, model cached
in a shared volume -- so the pt-PT voice is a new service, not a new
image.
- db.Backup + a `-backup` flag: VACUUM INTO, not a file copy. Petal runs
in WAL mode, so the newest committed pages may live in petal.db-wal;
copying the three files separately can capture a torn mid-checkpoint
state. VACUUM INTO reads one coherent snapshot without taking a write
lock, and emits a single file with no -wal/-shm companions. Refuses an
existing destination so a failed run can't destroy the last good
backup.
- deploy/backup-petal.sh: nightly snapshot, compress, push to millenia
over headscale with a post-transfer size check, prune both sides.
- deploy/petal.env.example: LLM_TIMEOUT raised 30s -> 90s for the
WAN+VPN round trip, since the voice and collocation passes send a
whole document and the timeout is a hard deadline on Complete.