Files
petal/deploy/README.md
T
prosolis 42d857a878 millenia: nightly encrypted backup, supervision, and hardened Piper units
The canonical instance -- the one with her actual writing -- turned out
to be the least protected thing in the estate:

- No scheduled backup at all; the newest snapshot was a month old. Now
  petal-backup.timer: VACUUM INTO, gzip, age-encrypt with the parodia
  public recipient, push to the VPS over headscale with a size check,
  prune both ends. Persistent=true because the box is not on 24/7.
  Neither machine can decrypt what it holds; the identity is offline.

- Petal ran as a bare ./petal with PPID 1, so a crash or reboot left it
  down until somebody noticed. Now petal.service, verified by kill -9.

- The Piper units retried forever without ever failing: RestartSec=3
  against systemd's default 10s window means the burst limit is never
  reached, which is how a dead service logged 26,800+ restarts over a
  day while read-aloud silently fell back to browser speech.
  StartLimitIntervalSec=300 makes a broken Piper show up in --failed.

backup-petal.sh now handles both deployment shapes (compose exec on the
VPS, local binary on millenia) and encrypts before anything leaves the
host. The VPS no longer uses it -- Petal rides parodia-backup there.
2026-07-27 06:35:34 -07:00

407 lines
16 KiB
Markdown

# Deploying Petal
Two deployments exist right now:
| | host | shape | status |
|---|---|---|---|
| **millenia** | `192.168.1.212` / `100.64.0.2` | bare binary in a screen session on `:8088`, Piper as user systemd units | **canonical** — her real writing lives here |
| **parodia** | `petal.parodia.dev` / `100.64.0.1` | docker compose behind the host's Traefik | staging; empty database |
millenia stays canonical until Phase 16's auth lands, so she only moves accounts
once. Everything below is the VPS side; the millenia Piper notes are kept in the
appendix because that instance still runs them.
---
## 1. The VPS stack
`docker-compose.yml` at the repo root brings up three containers:
- **petal** — the single Go binary with the frontend embedded. Publishes no host
port; Traefik is the only way in.
- **piper-en** / **piper-zh** — read-aloud. Each Piper HTTP server loads exactly
one voice, so English and Chinese are separate containers off one image, with
the models cached in a shared volume. They sit on an internal network with no
published ports, so only Petal can reach them. Adding pt-PT in Phase 21 is a
fourth service, not a new image.
They run as containers rather than the host systemd units millenia uses because
Piper was never actually installed on the VPS, and the `reala` account has no
lingering session to keep user units alive across logout.
### Prerequisites on the host
- Docker with the compose plugin, and the existing external `traefik` network
- A DNS A record for the hostname pointing at the VPS (`petal.parodia.dev` is
already in place)
### First deploy
```bash
ssh reala@100.64.0.1
git clone https://gitea.parodia.dev/drwily/petal.git ~/petal
cd ~/petal
cp deploy/petal.env.example .env
```
Then edit `.env`:
- `PETAL_UID` / `PETAL_GID``id -u` / `id -g` for this account. `./data` is a
bind mount, so the image's own `petal` user has no claim on it; a mismatch
shows up as `unable to open database file (14)` and a restart loop.
- `PETAL_BASIC_AUTH` — the interim edge gate, see §4.
- `LLM_MODEL` / `LLM_CHAT_MODEL` — see §3.
```bash
mkdir -p data/backups
docker compose up -d --build
docker compose ps # all three healthy
```
### Updating
```bash
cd ~/petal && git pull && docker compose up -d --build
```
The frontend is embedded in the binary, so a rebuild is the whole deploy. The
client polls `/api/version` (a hash of the built `index.html`) and offers a
refresh when it changes.
---
## 2. What Traefik does
Labels follow the convention the other services on this box use: the external
`traefik` network, the `web-secure` entrypoint, the `default` cert resolver and
`compression@file`. Petal adds its own response-header middleware
(`frame-ancestors 'self'`, HSTS, nosniff, `Referrer-Policy: same-origin`).
`/api/health` is deliberately on its own higher-priority router with no
middleware: a monitoring probe must not need a credential, and the endpoint
carries no user data.
---
## 3. The LLM link over headscale
The vLLM backend stays on millenia and is reached over headscale
(`100.64.0.2`). **This is the only cross-VPN dependency**, and by the
LLM-minimalism principle it never gates essential functionality — spell check,
gloss, vocabulary garden, search, export and read-aloud all keep working with
the link down, and the status bar shows the warm
`🌙 小助手在休息 · Petal's helper is resting · 文字已保存`.
`LLM_TIMEOUT` is raised from the local-network default of 30s to **90s**: the
voice and collocation passes send a whole document, the timeout is a hard
deadline on the completion call, and a WAN+VPN round trip eats the margin.
### How the link is exposed — a forwarder, not a rebind
vLLM stays bound to `127.0.0.1:8000`. `vllm-headscale-proxy.service` (a socat
unit, in this directory) adds a second listener on `100.64.0.2:8000` that
forwards to it.
The plan originally said to rebind vLLM itself. That turned out to be the
expensive option: `vllm-chat.service` is **shared** — Petal, Gogobee and Open
WebUI all point at `127.0.0.1:8000`, and Open WebUI stores its endpoint in its
own database rather than in env — so moving the bind address would mean editing
three consumers and reloading a 35B AWQ model, minutes of downtime for all of
them. The forwarder adds a door instead of moving one: local callers are
untouched, and the only new exposure is on the VPN interface.
It binds `100.64.0.2` specifically, **never** `0.0.0.0`: the far end of this
link is a public host, and the LAN has no business seeing an unauthenticated
inference endpoint.
```bash
sudo install -m 0644 deploy/vllm-headscale-proxy.service /etc/systemd/system/
sudo systemctl daemon-reload && sudo systemctl enable --now vllm-headscale-proxy
ss -lntp | grep 8000 # expect BOTH 127.0.0.1:8000 and 100.64.0.2:8000
```
The model id (`qwen3.6-35b`) goes into `LLM_MODEL` / `LLM_CHAT_MODEL` on the
VPS. Verified end to end: a grammar checkpoint from `petal.parodia.dev` returns
real suggestions in ~3s over the VPN.
---
## 4. Interim edge gate (delete when Phase 16 lands)
Petal authenticates nobody yet — `StaticResolver` hands every request the same
`local` user. On a public host that means anyone who finds the hostname can read
and write documents and fill the disk with image uploads, so Traefik holds the
door with basic auth until the OIDC flow exists.
Generate a credential:
```bash
htpasswd -nbB petal 'your-password' # or any bcrypt htpasswd generator
```
and put the resulting `user:hash` pair in `.env` as `PETAL_BASIC_AUTH`.
When Phase 16 lands, delete the `petal-auth` middleware label, the
`petal-health` router labels, and this section.
---
## 5. Backups
### On the VPS — folded into `parodia-backup`
Petal rides the host's existing offsite job (`/usr/local/bin/parodia-backup`,
`parodia-backup.timer`, nightly ~03:40): age-encrypted to S3, 14-day retention,
dead-man snitch. The host holds only the age *public* recipient, so it writes
backups it cannot itself decrypt.
```
push petal.db.age sqlite_file_dump /home/reala/petal/data/petal.db
```
`/home/reala/petal/.env` is in the same job's secrets tarball — it carries the
interim basic-auth hash and, from Phase 16, the OIDC client secret.
**Why `sqlite_file_dump` and not the script's existing `sqlite_dump`:** that
helper uses Python's `iterdump`, which **does not reproduce an FTS5 virtual
table**. It emits `documents_fts` as a raw `sqlite_master` row plus its shadow
tables, and replaying the result dies with `no such table: documents_fts`
verified by round-tripping a real dump on 2026-07-27. Petal's cross-document
search would have been silently missing after any restore. `sqlite_file_dump`
runs `VACUUM INTO` instead: a genuine database file, virtual tables intact, WAL
folded in, no write lock. Restore is a copy rather than a replay.
> If `apply.db` ever gains a virtual table, it needs the same treatment.
### Restore (VPS)
```bash
age -d -i <offline-identity> petal.db.age > /tmp/petal.db # from S3
cd ~/petal
docker compose stop petal # stop writers first
mv data/petal.db data/petal.db.before-restore # keep the current state
rm -f data/petal.db-wal data/petal.db-shm # a stale WAL against a new file
cp /tmp/petal.db data/petal.db
docker compose start petal
docker compose logs petal --tail 5 # expect "database ready"
```
To sanity-check an archive before committing to it, have Petal open it in a
scratch directory — a clean exit means it reads end to end:
```bash
mkdir -p /tmp/restore-check && cp /tmp/petal.db /tmp/restore-check/petal.db
docker run --rm -v /tmp/restore-check:/data --user "$(id -u):$(id -g)" \
--entrypoint sh petal:local -c '/app/petal -backup /data/verify.db'
```
### On millenia — `petal-backup.timer`
Until 2026-07-27 her actual writing had **no scheduled backup at all**; the
newest snapshot was a month old. It now runs nightly at 03:20
(`Persistent=true`, because the box isn't on 24/7 and a missed window would
otherwise be skipped silently):
```bash
sudo install -m 0644 deploy/petal-backup.service deploy/petal-backup.timer /etc/systemd/system/
sudo systemctl daemon-reload && sudo systemctl enable --now petal-backup.timer
sudo systemctl start petal-backup.service # prove it before trusting it
```
`deploy/backup-petal.sh` snapshots via `petal -backup`, gzips, **age-encrypts
with the parodia public recipient**, pushes to the VPS over headscale with a
post-transfer size check, and prunes both ends. The private identity is offline,
so neither millenia nor the VPS can decrypt what it is holding — verified.
A manual snapshot any time, no tooling required:
```bash
cd ~/petal && ./petal -backup ~/petal/backups/manual-$(date -u +%Y%m%dT%H%M%SZ).db
```
Restore is a copy — stop Petal, drop the file in as `data/petal.db`, remove any
stale `-wal`/`-shm`, start.
---
## 6. Encryption at rest
### VPS — `/home/reala/petal/data` is a LUKS volume
`deploy/setup-encrypted-data.sh` puts the data directory on LUKS2 over a sparse
file at `/var/lib/petal-crypt.img`. That covers `petal.db`, uploaded `images/`,
**and the TTS cache** — which is synthesized audio of her sentences and is easy
to forget.
LUKS-on-a-file rather than gocryptfs because Petal is SQLite in WAL mode: WAL
needs a shared-memory index (`-shm`) mapped consistently across processes, and
FUSE has a long history of subtle mmap/locking differences. A block device with
ext4 behaves exactly like a disk to SQLite, which is the only guarantee worth
having under a database.
**What it protects, honestly.** The key lives at `/etc/petal/dataset.key` on the
same host so the volume auto-unlocks at boot. That is a deliberate availability
tradeoff:
| | |
|---|---|
| protects against | a decommissioned or resold disk; reading the raw block device; casual browsing of a filesystem snapshot that excludes `/etc` |
| does **not** protect against | anyone holding the whole VM image — they get the keyfile with the ciphertext; or anything at all while the host is running and mounted |
Real protection from a provider-side snapshot needs the key off-box (fetched
over the VPN at boot). Considered, not chosen.
Two things this setup got wrong the first time, both caught by rehearsing a
reboot rather than trusting a clean run — worth knowing if you rebuild it:
- **Mounting over a directory hides its contents, it does not remove them.** The
first pass left the original plaintext `petal.db` and WAL sitting on the
unencrypted root filesystem, invisible under the mount. The script now shreds
the originals before mounting and refuses to continue if the mountpoint will
not come up empty.
- **`systemd-cryptsetup` was not installed**, so `/etc/crypttab` was ignored
entirely and the volume would never have unlocked at boot. The script now
refuses to run without the generator present.
Check it any time:
```bash
sudo ./deploy/setup-encrypted-data.sh --status
```
### The mount-liveness guard
The mountpoint directory exists whether or not the volume is mounted, so a boot
where the unlock failed would start Petal against an empty unencrypted
directory and quietly serve a blank database — the failure that looks like data
loss. `data/.volume-ok` lives on the encrypted filesystem and is bind-mounted
with `create_host_path: false`, turning that into a loud container start
failure:
```
Error response from daemon: invalid mount config for type "bind":
bind source path does not exist: /home/reala/petal/data/.volume-ok
```
Verified by unmounting and attempting a start.
**A true reboot has not been tested** — the VPS also runs matrix, lemmy, akkoma,
gitea and authentik, so rebooting it is your call. The boot path was rehearsed
through `local-fs.target`, which pulls the mount, which pulls the unlock.
### millenia is not encrypted at rest
LVM, no LUKS. Her canonical writing sits in plaintext on the home box. Backups
leaving it are age-encrypted; the disk itself is not.
---
## 7. Supervision and monitoring
Petal on millenia ran for months as a bare `./petal` with PPID 1 — no unit, no
screen session — so a crash or reboot left it down until someone noticed. It is
now `petal.service`:
```bash
sudo install -m 0644 deploy/petal.service /etc/systemd/system/
sudo systemctl daemon-reload && sudo systemctl enable --now petal.service
```
Verified by `kill -9`-ing it and watching systemd bring it back.
**Piper's silent-failure mode is fixed.** Both units now set
`StartLimitIntervalSec=300` / `StartLimitBurst=5`. With `RestartSec=3` and
systemd's default 10-second window, only ~3 restarts ever landed inside it, so
the burst limit was never reached and a dead service looped **26,800+ times over
a day without ever entering `failed`**. A genuinely broken Piper now shows up in
`systemctl --user --failed`.
### Still to do — an external probe
Nothing yet watches millenia from outside. uptime-kuma already runs on the VPS
and can reach millenia over headscale, so the missing piece is two monitors
(they need the uptime-kuma UI, hence not scripted here):
- `http://100.64.0.2:8088/api/health` — Petal itself
- a POST to `http://100.64.0.2:8088/api/tts` — catches a dead Piper, which
`/api/health` will not, because read-aloud degrades silently to browser
speech
---
## Appendix — Piper on millenia (user systemd units)
millenia still runs Piper as user services; these are the original notes.
```bash
scp deploy/piper.service deploy/setup-piper.sh 192.168.1.212:/tmp/
ssh 192.168.1.212 'cd /tmp && sudo ./setup-piper.sh'
```
`setup-piper.sh` creates `~/piper/venv`, installs `piper-tts[http]`, downloads
`en_US-amy-medium` into `~/piper/voices`, installs and enables `piper.service`
(loopback `:5005`), and smoke-tests it. Idempotent. Check it with
`systemctl status piper` / `journalctl -u piper -f`.
Chinese runs as a second instance (`piper-zh.service`, `:5006`,
`zh_CN-huayan-medium`):
```bash
scp deploy/piper-zh.service 192.168.1.212:~/.config/systemd/user/
ssh 192.168.1.212 'export XDG_RUNTIME_DIR=/run/user/$(id -u)
~/piper/venv/bin/python -m piper.download_voices zh_CN-huayan-medium --data-dir ~/piper/voices
systemctl --user daemon-reload && systemctl --user enable --now piper-zh.service'
```
Petal's env then carries `TTS_ENDPOINT=http://127.0.0.1:5005`,
`TTS_ENDPOINT_ZH=http://127.0.0.1:5006` and the matching voice ids. The handler
maps language → instance from config, so another language is another instance
plus an env pair, no code change.
**Piper version note:** piper-tts moved synthesis from `POST /` to
`POST /synthesize` in 1.6.0, with an identical request body. `TTS_PATH` selects
which — it defaults to `/`, and both the VPS compose and millenia's `start.sh`
now set `/synthesize`. If read-aloud starts returning 502 after a Piper upgrade,
that flag is the fix.
**The venv is fragile across Python upgrades.** On 2026-07-27 millenia's Piper
was found dead with **26,800+ failed restarts**, silently since the Jul 26
reboot — read-aloud had been falling back to browser Web Speech the whole time.
Root cause: an OS upgrade moved `/usr/bin/python3` from 3.13 to 3.14, and
`venv/bin/python3` is a *symlink to the system interpreter*, so the venv's
`lib/python3.13/site-packages` became invisible — `sys.path` contained no
site-packages at all. The failure surfaced as the misleading
`No module named piper.http_server` even though `http_server.py` was sitting
right there on disk.
Fix (what was done — recreating the venv, not repairing it):
```bash
systemctl --user stop piper.service piper-zh.service
mv ~/piper/venv ~/piper/venv.broken-py313
python3 -m venv ~/piper/venv
~/piper/venv/bin/pip install "piper-tts[http]"
~/piper/venv/bin/python -c 'import piper.http_server' # must not raise
systemctl --user start piper.service piper-zh.service
```
That reinstall lands 1.6.0, so it must be paired with `TTS_PATH=/synthesize` in
`start.sh` and a binary new enough to read that variable. Voices in
`~/piper/voices` survive and do not need re-downloading.
Worth knowing: `Restart=on-failure` will retry forever without ever alerting.
Neither service reports its health anywhere, which is why this went unnoticed
for a day. A `/api/tts` probe in uptime-kuma would have caught it.
Verify end to end (through Petal, including the ffmpeg transcode):
```bash
curl -sf -X POST localhost:8088/api/tts \
-H 'Content-Type: application/json' \
-d '{"text":"hello there","lang":"en-US"}' -o /tmp/petal-tts.mp3 \
&& file /tmp/petal-tts.mp3 # expect: MPEG ADTS, layer III
```
A second identical call is served from the cache; a language with no configured
instance returns 404 so the client falls back to Web Speech.