From 316b6b305d323a6d834d51a7228ff11f711b990a Mon Sep 17 00:00:00 2001 From: prosolis <5590409+prosolis@users.noreply.github.com> Date: Sun, 26 Jul 2026 21:48:51 -0700 Subject: [PATCH] Plan: DreamDict answers the Portuguese dictionary question MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit OPEN #6 assumed the pt-PT gloss was gated on finding a dataset of ECDICT's quality. It isn't — dreamdict already covers en, fr, pt-PT and zh, and its API maps almost 1:1 onto lexicon.Result. Replaces the data question with an integration one (OPEN #6a): HTTP service, build-time extraction into Petal's embedded gz format, or importing the dictionary package and opening dict.db read-only. Argues for the third — same CGO-free sqlite driver Petal already depends on, no runtime service, and it deletes ~11.6 MB of embedded data plus the ECDICT build scripts. Adds the migration caution that matters most: the zh path is in daily use, so wire pt-PT/fr first (nothing to regress) and leave zh on ECDICT until CC-CEDICT gloss quality has been compared on her real lookups. --- MULTIUSER_PLAN.md | 110 +++++++++++++++++++++++++++++++++++++--------- 1 file changed, 90 insertions(+), 20 deletions(-) diff --git a/MULTIUSER_PLAN.md b/MULTIUSER_PLAN.md index 3f98ddd..c2b52b3 100644 --- a/MULTIUSER_PLAN.md +++ b/MULTIUSER_PLAN.md @@ -210,7 +210,7 @@ dataset identical for everyone. --- -## 6. Phase D — per-user language +## 6. Phase D — per-user language (and DreamDict) Tracked here because it lands on the same `users` row and shouldn't be designed twice. @@ -219,26 +219,89 @@ twice. language — the one glosses and explanations are written in. Mandarin ships today; European Portuguese (pt-PT, explicitly not pt-BR) is wanted; French is possible. -That makes native language a `users` column and turns these into per-user -lookups: +### OPEN #6 is answered: DreamDict -| Piece | Today | Notes | -|---|---|---| -| Gloss / lexicon data | ECDICT, English↔Chinese | pt-PT needs its own source — **this is the hard part**, not the code | -| LLM prompt copy | Mandarin-first bilingual | `internal/llm/prompts.go` | -| Companion tips | bilingual `tips.ts` | per-language copy | -| TTS L1 voice | Piper zh_CN-huayan | Piper has pt-PT voices | -| Font stacks | CJK fallbacks | not needed for Latin-script L1 | +The original worry here was data sourcing — Petal's gloss comes from ECDICT +(English↔Chinese), and a pt-PT equivalent of comparable quality and license +looked like the blocker. -English-side machinery (nspell en-US, the IPA/phonetic dataset, the EN Piper -voice) is unaffected and stays shared. +`~/git/dreamdict` already solves it, and more completely than expected. It +supports **en, fr, pt-PT, and zh** (~136k/56k/136k/121k words), and its shape maps +almost 1:1 onto `lexicon.Result`: -**OPEN #6:** the gloss dataset is the blocker, not the plumbing. Is there an -open English↔Portuguese dictionary of ECDICT's quality and license? If not, this -phase is gated on data sourcing and the code work is comparatively trivial. Worth -answering before scheduling it. +| Petal field | DreamDict | +|---|---| +| `Gloss` | `Translate(word, "en", L1)` | +| `Phonetic` | pronunciation (CMU + IPA for en, Wiktionary IPA elsewhere) | +| `Definitions` | `Define(word, lang)` — curated sources ranked above Wiktionary | +| `Synonyms` | `Synonyms(word, lang)` | ---- +It also carries data Petal has no equivalent for and could use: `Antonyms`, +`Frequency`, `Difficulty`, and `Etymology`. + +So Phase D stops being gated on data and becomes an integration decision. + +### OPEN #6a (new): how to integrate + +**Option 1 — HTTP client.** Petal calls DreamDict on localhost:7777, exactly the +pattern already used for Piper TTS (including graceful degradation when it's +down). +- *For:* zero coupling, DreamDict updates independently, all endpoints available. +- *Against:* a second service Petal now depends on at runtime, and the gloss is a + 350ms hover tooltip where "the dictionary service is down" is a visible + regression from today's always-there embedded data. + +**Option 2 — build-time extraction.** A script (sibling to the existing +`scripts/build_gloss.py`) generates Petal's embedded `.json.gz` datasets per +language from DreamDict's `dict.db`. +- *For:* preserves the embedded/offline property exactly; no runtime dependency; + no architectural change at all. +- *Against:* every language multiplies the binary (the four current gz files are + already ~11.6 MB); updating the dictionary means rebuilding and redeploying + Petal; the richer fields are lost unless separately extracted. + +**Option 3 — import the package, open `dict.db` read-only.** DreamDict's +`internal/dictionary` is a plain library with `NewReadOnly(dbPath)`, and its only +dependency is `modernc.org/sqlite` — the same CGO-free driver Petal already uses. +Petal opens `dict.db` as a second read-only handle beside `petal.db`. +- *For:* no service, no HTTP, no new dependency, lookups stay local-file fast, + all four languages at once, and it deletes ~11.6 MB of embedded gz plus the + ECDICT build scripts. One dictionary, maintained once, shared with GogoBee. +- *Against:* Petal stops being a self-contained binary in the "just run it" sense + — `dict.db` has to be deployed alongside. In practice Petal already ships a + data directory (`petal.db`, images, TTS cache), so this is a smaller loss than + it first sounds. + +**My recommendation: Option 3.** It is the only one that gets all four languages, +keeps lookups offline and instant, and *removes* code rather than adding a +subsystem. Option 1's runtime dependency buys flexibility Petal doesn't need for +a dictionary that changes a few times a year. + +**Prerequisite:** DreamDict's module path is currently `module dreamdict`, which +isn't fetchable. Importing it needs the module renamed to something like +`gitea.parodia.dev/drwily/dreamdict` (or a local `replace` directive for +development). Small, but it must happen first. + +### Migration caution + +Whichever option wins, the zh path is **currently working and in daily use**. The +gloss quality difference between ECDICT and CC-CEDICT is unknown and matters more +than the architecture. + +Proposal: introduce DreamDict behind Petal's existing lexicon interface as a +*provider*, wire pt-PT and fr to it first (nothing to regress — they don't exist +yet), and keep zh on ECDICT until the two have been compared on real lookups from +her actual documents. Converge only if quality holds. This also de-risks the whole +change: if DreamDict turns out to be a poor fit, only the unshipped languages are +affected. + +### Still per-user regardless + +Native language becomes a `users` column, and these become per-user lookups: +LLM prompt copy (`internal/llm/prompts.go`, currently Mandarin-first), companion +tips (`tips.ts`), the L1 Piper voice (Piper has pt-PT voices), and the CJK font +stacks (not needed for Latin-script L1). English-side machinery — nspell en-US, +the phonetic dataset, the EN voice — is unaffected and stays shared. ## 7. Suggested sequence @@ -250,7 +313,10 @@ answering before scheduling it. 4. Image store table + migration (same phase, per OPEN #5). 5. Provision the second real account; migrate `local`'s data. 6. `localStorage` namespacing. -7. Per-user language, gated on the dataset question. +7. Per-user language. **No longer gated on data** — DreamDict covers all four + languages. Sequence within it: rename DreamDict's module path → wire it in as + a lexicon provider → pt-PT/fr first → compare zh quality → converge if it + holds. --- @@ -279,7 +345,11 @@ answering before scheduling it. 4. Migration script vs. admin endpoint for moving `local`'s data? (OPEN #4) 5. Fix the image store alongside auth, or ship auth with it as a known limitation? (OPEN #5) -6. Is there a usable open English↔European-Portuguese dictionary dataset? - (OPEN #6 — gates Phase D entirely) +6. ~~Is there a usable open English↔European-Portuguese dictionary dataset?~~ + **Answered: DreamDict**, which covers en/fr/pt-PT/zh. The live question is now + *how* to integrate it — HTTP service, build-time extraction, or importing the + package and opening `dict.db` read-only. (OPEN #6a; I recommend the third) +7. Does replacing the zh gloss (ECDICT → CC-CEDICT) risk regressing a feature in + daily use, and should zh stay on ECDICT until the two are compared? Anything above that reads as settled but shouldn't be is also fair game.