Give read-aloud a Portuguese voice, and a slower one

Phase 21's infra half. Two things the pt-PT pair needs from TTS, and one
thing every learner has wanted since Phase 11.

**A language is no longer a code change.** The handler knew exactly two
languages, named in the Config struct: English on TTS_ENDPOINT and Chinese
on TTS_ENDPOINT_ZH. Petal now discovers its Piper instances from the
environment — English keeps the unsuffixed pair it has always had, and
every other language is a TTS_ENDPOINT_<LANG>/TTS_VOICE_<LANG> pair — so
fr and es cost a compose service and two lines of .env. <LANG> is the base
tag, because an environment variable name cannot hold pt-PT's hyphen and
only one Portuguese model is loaded either way. A language configured by
halves is dropped rather than routed: half a configuration should reach
the client as "no voice here, use Web Speech", not as an instance that
errors on every tap. The startup line now names the voices it actually
resolved rather than the English endpoint it was handed — the same lesson
the dictionary line learned last week.

**pt_PT-tugão-medium is the only European voice Piper ships.** The other
five pt models in the catalogue are Brazilian, so the default anyone
reaches for is the wrong country — the same trap as `dictionary-pt`
packaging VERO, arriving through the catalogue rather than through the
model. Named explicitly in compose, with the query that checks it in the
deploy README.

**The slow replay** (SUGGESTIONS §5e) is `slow: true` on /api/tts, raising
Piper's length_scale to ~4/3. Piper stretches durations rather than
resampling, so it stays a voice instead of a groan. The pace is part of
the cache key — without it the slow replay of a word already heard at
normal speed would be served back at normal speed, which is the one
request where the difference is the whole point. 🐢 sits beside 🔊 on the
word card, the selection bubble and the garden flashcard; the Web Speech
fallback slows too, so the button means the same thing when Piper is down.

**And the other reading gets her own voice.** The `alsoIn` block — the
Portuguese sense of a word that is also English — now speaks in the pair's
locale, which the pack names (`locale`) rather than anything inferring it
from the letters. "comum" is spelled identically in both halves; a
detector would have to guess, and this is the same reason the gloss shows
both directions instead of picking one.

Tests: config discovery (both existing deployment shapes, half-configured
languages dropped, the pre-map voice defaults preserved), the slow scale
and its separate cache entry, pt routing on the base tag with pt-BR
landing on the European instance, and speech.ts's request body. The i18n
shape suite now asserts every pack names a speakable locale in its own
language — and that pt-PT's is not pt-BR.

Verified: go build/vet/test, tsc, vitest 125/125, vite build. Live smoke
against two fake Piper servers: en/pt × normal/slow all reached the right
instance at the right length_scale with four distinct cache entries, and
an unconfigured language still 404s.
This commit is contained in:
prosolis
2026-07-27 13:21:45 -07:00
parent ccb43e5a4d
commit 24c3533e18
18 changed files with 595 additions and 66 deletions
+81 -12
View File
@@ -2,6 +2,7 @@ package config
import (
"os"
"strings"
"time"
)
@@ -29,10 +30,16 @@ type Config struct {
// TTS (read-aloud). Off unless TTSEndpoint is set — when empty, the /api/tts
// route isn't mounted and the frontend falls back to the browser's Web Speech
// API. Endpoint points at a local Piper HTTP server.
TTSEndpoint string // Piper instance serving the English voice
TTSEndpointZH string // Piper instance serving the Chinese voice; empty = zh falls back to Web Speech
TTSVoiceEN string // Piper voice id for English (e.g. en_US-amy-medium)
TTSVoiceZH string // Piper voice id for Chinese (e.g. zh_CN-huayan-medium)
TTSEndpoint string // Piper instance serving the English voice; also the on/off switch
// TTSVoices is every language Petal can read aloud, keyed by base language
// tag ("en", "zh", "pt", …). Each Piper server loads exactly one model, so
// a language *is* an instance — and the instances are discovered from the
// environment rather than named in this struct: one
// TTS_ENDPOINT_<LANG>/TTS_VOICE_<LANG> pair per language, so the fr and es
// pairs cost a compose service and two lines of .env rather than a code
// change. English keeps the unsuffixed TTS_ENDPOINT/TTS_VOICE_EN it has
// always had.
TTSVoices map[string]TTSVoice
// TTSPath is the path Piper serves synthesis on. Piper moved it from "/" to
// "/synthesize" in 1.6.0 with an unchanged request body, so this is a
// version knob, not a feature: millenia's older server keeps the default,
@@ -56,6 +63,12 @@ type Config struct {
AllowedSubs string
}
// TTSVoice is one Piper instance and the single voice it has loaded.
type TTSVoice struct {
Endpoint string
Voice string
}
// AuthEnabled reports whether real logins are configured. When false, Petal
// resolves every request to the local user.
func (c *Config) AuthEnabled() bool {
@@ -77,14 +90,12 @@ func Load() *Config {
LLMChatModel: env("LLM_CHAT_MODEL", ""),
LLMTimeout: envDuration("LLM_TIMEOUT", 30*time.Second),
TTSEndpoint: env("TTS_ENDPOINT", ""),
TTSEndpointZH: env("TTS_ENDPOINT_ZH", ""),
TTSVoiceEN: env("TTS_VOICE_EN", "en_US-amy-medium"),
TTSVoiceZH: env("TTS_VOICE_ZH", "zh_CN-huayan-medium"),
TTSPath: env("TTS_PATH", "/"),
TTSCacheDir: env("TTS_CACHE_DIR", "./data/tts"),
TTSTimeout: envDuration("TTS_TIMEOUT", 15*time.Second),
TTSFormat: env("TTS_AUDIO_FORMAT", "mp3"),
TTSEndpoint: env("TTS_ENDPOINT", ""),
TTSVoices: ttsVoices(os.Environ()),
TTSPath: env("TTS_PATH", "/"),
TTSCacheDir: env("TTS_CACHE_DIR", "./data/tts"),
TTSTimeout: envDuration("TTS_TIMEOUT", 15*time.Second),
TTSFormat: env("TTS_AUDIO_FORMAT", "mp3"),
AuthentikURL: env("AUTHENTIK_URL", ""),
AuthentikClientID: env("AUTHENTIK_CLIENT_ID", ""),
@@ -93,6 +104,64 @@ func Load() *Config {
}
}
// ttsVoices reads the Piper instances out of an environment slice (as returned
// by os.Environ) into a map keyed by base language tag.
//
// English is the unsuffixed pair, TTS_ENDPOINT + TTS_VOICE_EN, because that is
// what every deployment already sets and read-aloud has always been English
// first. Every other language is a TTS_ENDPOINT_<LANG>/TTS_VOICE_<LANG> pair,
// discovered rather than enumerated — TTS_ENDPOINT_ZH is what millenia and the
// VPS already use, and TTS_ENDPOINT_PT is all the Portuguese pair needs.
//
// <LANG> is the *base* tag: an environment variable name cannot hold the hyphen
// in "pt-PT", and the handler routes on the base tag anyway (a request for
// pt-PT, pt-BR or bare pt reaches the same instance, because there is only one
// Portuguese voice loaded). A pair is ignored unless both halves are set: half
// a configuration should read as "no voice for this language" and fall back to
// the browser, not as an instance that answers every request with an error.
func ttsVoices(environ []string) map[string]TTSVoice {
vals := make(map[string]string, len(environ))
for _, kv := range environ {
if k, v, ok := strings.Cut(kv, "="); ok {
vals[k] = v
}
}
voices := map[string]TTSVoice{}
add := func(lang, endpoint, voice string) {
endpoint = strings.TrimRight(strings.TrimSpace(endpoint), "/")
voice = strings.TrimSpace(voice)
if endpoint == "" || voice == "" {
return
}
voices[lang] = TTSVoice{Endpoint: endpoint, Voice: voice}
}
// The two languages that shipped before this was a map keep their voice
// defaults, so an existing deployment that names only the endpoints (as
// millenia's unit does) sounds exactly as it did.
voiceOr := func(key, fallback string) string {
if v := strings.TrimSpace(vals[key]); v != "" {
return v
}
return fallback
}
add("en", vals["TTS_ENDPOINT"], voiceOr("TTS_VOICE_EN", "en_US-amy-medium"))
for k, endpoint := range vals {
suffix, ok := strings.CutPrefix(k, "TTS_ENDPOINT_")
if !ok || suffix == "" {
continue
}
voice := vals["TTS_VOICE_"+suffix]
if suffix == "ZH" {
voice = voiceOr("TTS_VOICE_ZH", "zh_CN-huayan-medium")
}
add(strings.ToLower(suffix), endpoint, voice)
}
return voices
}
func env(key, fallback string) string {
if v := os.Getenv(key); v != "" {
return v
+83
View File
@@ -0,0 +1,83 @@
package config
import "testing"
// The Piper instances are discovered from the environment rather than named in
// code, so that a new pair costs a compose service and two .env lines. These
// assert the discovery rule, including the two shapes that already exist in the
// wild (millenia's systemd unit and the VPS compose file).
func TestTTSVoicesDiscovery(t *testing.T) {
voices := ttsVoices([]string{
"TTS_ENDPOINT=http://piper-en:5000",
"TTS_VOICE_EN=en_US-amy-medium",
"TTS_ENDPOINT_ZH=http://piper-zh:5000/",
"TTS_VOICE_ZH=zh_CN-huayan-medium",
"TTS_ENDPOINT_PT=http://piper-pt:5000",
"TTS_VOICE_PT=pt_PT-tugão-medium",
// Noise that must not become a language.
"TTS_PATH=/synthesize",
"PATH=/usr/bin",
})
want := map[string]TTSVoice{
"en": {"http://piper-en:5000", "en_US-amy-medium"},
// The trailing slash is trimmed here so the synthesis path concatenates
// cleanly rather than producing a double slash at every call site.
"zh": {"http://piper-zh:5000", "zh_CN-huayan-medium"},
"pt": {"http://piper-pt:5000", "pt_PT-tugão-medium"},
}
if len(voices) != len(want) {
t.Fatalf("discovered %v, want %v", voices, want)
}
for lang, w := range want {
if voices[lang] != w {
t.Errorf("%s = %+v, want %+v", lang, voices[lang], w)
}
}
}
// Half a configuration is not a language. An endpoint with no voice (or the
// reverse) must read as "no voice for this language" — a 404 the client answers
// by falling back to Web Speech — rather than as an instance that exists and
// errors on every request.
func TestTTSVoicesIgnoresHalfConfiguredLanguages(t *testing.T) {
voices := ttsVoices([]string{
"TTS_ENDPOINT=http://piper-en:5000",
"TTS_VOICE_EN=en_US-amy-medium",
"TTS_ENDPOINT_FR=http://piper-fr:5000", // no TTS_VOICE_FR
"TTS_VOICE_ES=es_ES-davefx-medium", // no TTS_ENDPOINT_ES
})
if _, ok := voices["fr"]; ok {
t.Errorf("fr routed with no voice configured")
}
if _, ok := voices["es"]; ok {
t.Errorf("es routed with no endpoint configured")
}
if len(voices) != 1 {
t.Errorf("discovered %v, want English only", voices)
}
}
// A deployment that predates the map names only the endpoints and relies on the
// voice defaults; it must sound exactly as it did.
func TestTTSVoicesKeepsTheOriginalDefaults(t *testing.T) {
voices := ttsVoices([]string{
"TTS_ENDPOINT=http://127.0.0.1:5005",
"TTS_ENDPOINT_ZH=http://127.0.0.1:5006",
})
if got := voices["en"].Voice; got != "en_US-amy-medium" {
t.Errorf("en voice = %q, want the default", got)
}
if got := voices["zh"].Voice; got != "zh_CN-huayan-medium" {
t.Errorf("zh voice = %q, want the default", got)
}
}
// Read-aloud is off when no English instance is configured; nothing else may
// switch it on. (tts.New gates on TTSEndpoint, so a stray TTS_ENDPOINT_PT with
// no English sibling must not produce a routable map that outlives that gate.)
func TestTTSVoicesEmptyWithoutEndpoints(t *testing.T) {
if voices := ttsVoices([]string{"TTS_VOICE_EN=en_US-amy-medium"}); len(voices) != 0 {
t.Errorf("discovered %v, want none", voices)
}
}
+45 -17
View File
@@ -19,6 +19,7 @@ import (
"os"
"os/exec"
"path/filepath"
"sort"
"strings"
"time"
"unicode/utf8"
@@ -36,6 +37,12 @@ const maxTextBytes = 4000
// with the words, mirroring the old utterance.rate = 0.95. Higher = slower.
const lengthScale = 1.1
// slowLengthScale is the "say it slower" replay (SUGGESTIONS §5e): roughly 0.75×
// the normal pace, which is the speed listening drills have used for decades.
// Piper stretches durations rather than resampling, so the voice keeps its pitch
// instead of turning into a slowed tape.
const slowLengthScale = lengthScale / 0.75
// audioFormat describes one output encoding: the cache-file extension, the
// response Content-Type, and the ffmpeg args that turn Piper's WAV (on stdin)
// into this format (on stdout). A nil ffmpegArgs means "serve the WAV as-is".
@@ -105,16 +112,14 @@ func New(cfg *config.Config) (*Handler, bool) {
format = formats["mp3"]
}
// Map by base language so en-US, en-GB, etc. all resolve to the English
// instance (the client sends BCP-47 tags like the old Web Speech path did).
// A language is only routable when both its endpoint and voice are set;
// otherwise the client falls back to Web Speech for that language.
// Keyed by base language so en-US, en-GB — and pt-PT, pt-BR, bare pt —
// resolve to the one instance that has that language's model loaded (the
// client sends BCP-47 tags, as the old Web Speech path did). Config has
// already dropped any language configured by halves, so an unroutable
// language reaches the client as a 404 and falls back to Web Speech.
routes := map[string]route{}
if cfg.TTSVoiceEN != "" {
routes["en"] = route{strings.TrimRight(cfg.TTSEndpoint, "/"), cfg.TTSVoiceEN}
}
if cfg.TTSEndpointZH != "" && cfg.TTSVoiceZH != "" {
routes["zh"] = route{strings.TrimRight(cfg.TTSEndpointZH, "/"), cfg.TTSVoiceZH}
for lang, v := range cfg.TTSVoices {
routes[lang] = route{endpoint: v.Endpoint, voice: v.Voice}
}
if err := os.MkdirAll(cfg.TTSCacheDir, 0o755); err != nil {
@@ -132,6 +137,19 @@ func New(cfg *config.Config) (*Handler, bool) {
}, true
}
// Languages lists the base language tags this handler can synthesize, sorted,
// each with the voice serving it — for the startup line, so a deployment says
// which sidecars it actually reached rather than which ones it was configured
// to want.
func (h *Handler) Languages() []string {
out := make([]string, 0, len(h.routes))
for lang, rt := range h.routes {
out = append(out, lang+"="+rt.voice)
}
sort.Strings(out)
return out
}
// Routes mounts the synthesis endpoint. Mount under "/tts" so the full path is
// POST /api/tts.
func (h *Handler) Routes() chi.Router {
@@ -140,11 +158,13 @@ func (h *Handler) Routes() chi.Router {
return r
}
// synthRequest is the body the editor posts: a passage and the BCP-47 language
// tag it's written in (e.g. "en-US", "zh-CN").
// synthRequest is the body the editor posts: a passage, the BCP-47 language tag
// it's written in (e.g. "en-US", "zh-CN", "pt-PT"), and whether to say it slowly
// — the replay a learner reaches for when the sentence went past too fast.
type synthRequest struct {
Text string `json:"text"`
Lang string `json:"lang"`
Slow bool `json:"slow"`
}
// synth resolves a voice for the requested language, returns cached audio when
@@ -180,9 +200,17 @@ func (h *Handler) synth(w http.ResponseWriter, r *http.Request) {
return
}
// Content-addressed: identical (voice, text) → identical clip. The format
// extension keeps encodings from colliding in the same dir.
sum := sha256.Sum256([]byte(rt.voice + "\n" + text))
scale := lengthScale
if req.Slow {
scale = slowLengthScale
}
// Content-addressed: identical (voice, pace, text) → identical clip. The pace
// belongs in the key — without it the slow replay of a word already heard at
// normal speed would be served from cache at normal speed, which is the one
// request where the difference is the whole point. The format extension keeps
// encodings from colliding in the same dir.
sum := sha256.Sum256([]byte(fmt.Sprintf("%s\n%.3f\n%s", rt.voice, scale, text)))
name := hex.EncodeToString(sum[:])[:32] + h.format.ext
path := filepath.Join(h.cacheDir, name)
@@ -191,7 +219,7 @@ func (h *Handler) synth(w http.ResponseWriter, r *http.Request) {
return
}
audio, err := h.synthesize(r.Context(), rt, text)
audio, err := h.synthesize(r.Context(), rt, text, scale)
if err != nil {
http.Error(w, "synthesis failed", http.StatusBadGateway)
fmt.Fprintf(os.Stderr, "tts: synthesize: %v\n", err)
@@ -222,11 +250,11 @@ func (h *Handler) serve(w http.ResponseWriter, r *http.Request, path string) {
// synthesize POSTs to the route's Piper instance, then transcodes the returned
// WAV when the configured format calls for it.
func (h *Handler) synthesize(ctx context.Context, rt route, text string) ([]byte, error) {
func (h *Handler) synthesize(ctx context.Context, rt route, text string, scale float64) ([]byte, error) {
body, _ := json.Marshal(map[string]any{
"text": text,
"voice": rt.voice,
"length_scale": lengthScale,
"length_scale": scale,
})
httpReq, err := http.NewRequestWithContext(ctx, http.MethodPost, rt.endpoint+h.synthURI, bytes.NewReader(body))
if err != nil {
+91 -3
View File
@@ -22,6 +22,7 @@ func newStubPiper(t *testing.T, body []byte) (*httptest.Server, *int32, *synthEc
_ = json.NewDecoder(r.Body).Decode(&req)
last.voice, _ = req["voice"].(string)
last.text, _ = req["text"].(string)
last.scale, _ = req["length_scale"].(float64)
w.Header().Set("Content-Type", "audio/wav")
_, _ = w.Write(body)
}))
@@ -29,7 +30,10 @@ func newStubPiper(t *testing.T, body []byte) (*httptest.Server, *int32, *synthEc
return srv, &calls, last
}
type synthEcho struct{ voice, text string }
type synthEcho struct {
voice, text string
scale float64
}
// newHandler builds a wav-format handler (no ffmpeg) pointed at a stub server.
func newHandler(t *testing.T, endpoint string) *Handler {
@@ -88,7 +92,12 @@ func TestSynthPathNormalisation(t *testing.T) {
func post(t *testing.T, h *Handler, text, lang string) *httptest.ResponseRecorder {
t.Helper()
b, _ := json.Marshal(synthRequest{Text: text, Lang: lang})
return postReq(t, h, synthRequest{Text: text, Lang: lang})
}
func postReq(t *testing.T, h *Handler, body synthRequest) *httptest.ResponseRecorder {
t.Helper()
b, _ := json.Marshal(body)
req := httptest.NewRequest(http.MethodPost, "/", bytes.NewReader(b))
rr := httptest.NewRecorder()
h.synth(rr, req)
@@ -192,8 +201,87 @@ func TestTextIsCapped(t *testing.T) {
}
}
// The slow replay is the whole of SUGGESTIONS §5e: same text, same voice, more
// time per phoneme.
func TestSlowRequestStretchesTheVoice(t *testing.T) {
srv, _, last := newStubPiper(t, []byte("RIFF....fake-wav"))
h := newHandler(t, srv.URL)
if rr := postReq(t, h, synthRequest{Text: "reception", Lang: "en-US"}); rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200", rr.Code)
}
if last.scale != lengthScale {
t.Fatalf("normal length_scale = %v, want %v", last.scale, lengthScale)
}
if rr := postReq(t, h, synthRequest{Text: "reception", Lang: "en-US", Slow: true}); rr.Code != http.StatusOK {
t.Fatalf("slow status = %d, want 200", rr.Code)
}
if last.scale != slowLengthScale {
t.Fatalf("slow length_scale = %v, want %v", last.scale, slowLengthScale)
}
if slowLengthScale <= lengthScale {
t.Fatalf("slowLengthScale %v is not slower than %v", slowLengthScale, lengthScale)
}
}
// The pace has to be part of the cache key. Without it, asking for the slow
// replay of a word already heard at normal speed serves the normal clip — the
// one request where hearing the difference is the entire point.
func TestSlowClipIsNotServedFromTheNormalCache(t *testing.T) {
srv, calls, last := newStubPiper(t, []byte("RIFF....fake-wav"))
h := newHandler(t, srv.URL)
postReq(t, h, synthRequest{Text: "reception", Lang: "en-US"})
postReq(t, h, synthRequest{Text: "reception", Lang: "en-US", Slow: true})
if *calls != 2 {
t.Fatalf("piper calls = %d, want 2 (the slow clip is a different clip)", *calls)
}
if last.scale != slowLengthScale {
t.Fatalf("second call length_scale = %v, want the slow one", last.scale)
}
// …and each pace still caches on its own.
postReq(t, h, synthRequest{Text: "reception", Lang: "en-US", Slow: true})
postReq(t, h, synthRequest{Text: "reception", Lang: "en-US"})
if *calls != 2 {
t.Fatalf("piper calls = %d, want 2 (both paces now cached)", *calls)
}
}
// A Portuguese request must reach the Portuguese instance on the base tag alone:
// env var names cannot hold the hyphen in pt-PT, so config keys the map on "pt"
// and the handler has to meet it there. pt-BR resolves to the same instance
// because there is only one Portuguese voice loaded — and it is the European one.
func TestPortugueseRoutesOnTheBaseTag(t *testing.T) {
enSrv, enCalls, _ := newStubPiper(t, []byte("EN-wav"))
ptSrv, ptCalls, ptLast := newStubPiper(t, []byte("PT-wav"))
h := &Handler{
routes: map[string]route{
"en": {strings.TrimRight(enSrv.URL, "/"), "en_US-amy-medium"},
"pt": {strings.TrimRight(ptSrv.URL, "/"), "pt_PT-tugão-medium"},
},
cacheDir: t.TempDir(),
format: formats["wav"],
client: http.DefaultClient,
}
// Distinct text per tag, so a cache hit can't stand in for a route.
for i, tag := range []string{"pt-PT", "pt", "pt-BR"} {
if rr := post(t, h, strings.Repeat("receção ", i+1), tag); rr.Code != http.StatusOK {
t.Fatalf("%s status = %d, want 200", tag, rr.Code)
}
}
if *ptCalls != 3 || *enCalls != 0 {
t.Fatalf("calls en=%d pt=%d, want en=0 pt=3", *enCalls, *ptCalls)
}
if ptLast.voice != "pt_PT-tugão-medium" {
t.Fatalf("pt voice = %q, want the European Portuguese voice", ptLast.voice)
}
}
func TestBaseLang(t *testing.T) {
cases := map[string]string{"en-US": "en", "EN_gb": "en", "zh-CN": "zh", "en": "en", "": ""}
cases := map[string]string{"en-US": "en", "EN_gb": "en", "zh-CN": "zh", "pt-PT": "pt", "en": "en", "": ""}
for in, want := range cases {
if got := baseLang(in); got != want {
t.Errorf("baseLang(%q) = %q, want %q", in, got, want)