Files
petal/internal/lexicon/lexicon.go
T
prosolis ccb43e5a4d Phase 21: Petal learns to be an English+Portuguese pair
The plan said "Hunspell pt-PT vendored like en-US". Measuring that first is
what saved it: nspell expands affixes eagerly on construction, and European
Portuguese's 1,340 rules over 44,257 stems want over a gigabyte of browser
heap — ~340 MB for the first 12,000 entries, and no return at all after three
minutes on the whole file. So the expansion runs once at build time instead:
1,039,058 forms, 2.66 MB gzipped, read by the same nspell in 842 ms.

The obvious npm package would also have shipped the wrong language. Both
dictionary-pt and dictionary-pt-br carry VERO, the Brazilian word list, so
vendoring by name puts pt-BR spellings behind a pt-PT label — the drift
SUGGESTIONS §3 warns about, arriving through the packaging where no reviewer
can see it. The source is Projecto Natura's, and the build script now asserts
the fault lines (receção in, recepção out) before writing anything.

Spellcheck consults both dictionaries and flags only what both reject, which
is the no-detector answer to a pair with no script boundary. The word card
does the same in the other direction: "data" is a word in both languages, so
Petal shows both readings rather than guessing which she meant.

Writing the tests caught the one real bug — extendedAlphabet was a snapshot
while correct/suggest read live, and her dictionary arrives after English, so
every lookup would have resolved "cora" while the underlines were already
right.

Not done, and not claimed: the pack has not been read by a pt-PT speaker, and
the Piper voice is deferred with the deploy.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 12:43:02 -07:00

314 lines
10 KiB
Go

package lexicon
import (
"bytes"
"compress/gzip"
"encoding/json"
"fmt"
"io"
"strings"
"sync"
)
// Meaning is one sense of a word: its part of speech, the gloss, and an optional
// usage example. The frontend renders a few of these in the word popover.
type Meaning struct {
PartOfSpeech string `json:"part_of_speech"`
Definition string `json:"definition"`
Example string `json:"example,omitempty"`
}
// Result is the full lookup for one word. Either list may be empty (the word
// isn't a headword in that dataset); the frontend handles a partial or empty
// result gracefully. Gloss is the Chinese translation (empty when the word isn't
// in the gloss dataset) — shown first in the popover for the Mandarin-speaking
// writer.
type Result struct {
Word string `json:"word"`
Gloss string `json:"gloss"`
Phonetic string `json:"phonetic"` // IPA for the English word; "" when absent
Definitions []Meaning `json:"definitions"`
Synonyms []string `json:"synonyms"`
// The fields below only ever come from DreamDict; the embedded datasets
// leave them at their unknown values, and the popover hides them.
// Frequency is how common the word is (higher = more common). 0 means
// unknown, which is DreamDict's own convention — a word it carries but has
// no corpus count for is indistinguishable from a word it doesn't carry,
// and the popover treats both the same way.
Frequency int `json:"frequency"`
// Difficulty runs 0.0 (easiest) to 1.0 (hardest); -1 means unknown. It is a
// sentinel rather than an omitted field because 0.0 is a real, meaningful
// score and `omitempty` would erase it.
Difficulty float64 `json:"difficulty"`
// Etymology is free-form Wiktionary prose, trimmed to a line. Where the
// word came from is a real hook for a writer whose own language shares
// Latin roots with English — "ephemeral" is much easier to keep once you
// have seen efémero next to it.
Etymology string `json:"etymology"`
// Reverse is the same token read as a word of the writer's own language,
// present only when it is one. Absent for every writer whose pair is not
// Latin-script, and for the overwhelming majority of words in one that is.
Reverse *Reverse `json:"reverse,omitempty"`
}
// Reverse is a lookup in the other direction: the token treated as a word of the
// writer's language, translated into English.
//
// It exists because a Latin-script pair has no script boundary to tell the two
// halves apart. In English+Chinese, "which language is this word?" answers
// itself. In English+Portuguese it does not: *sale*, *casa*, *comum*, *tarde*
// and *ali* are all real words on both sides, and *chat* and *pain* are the
// French versions of the same trap.
//
// Petal does not guess. It asks both directions and shows whatever comes back,
// which needs no detector, cannot be wrong about someone's writing, and — for a
// learner — is more interesting than a correct guess would have been.
type Reverse struct {
// Lang is the language this reading is in, so the card can label it.
Lang string `json:"lang"`
// Gloss is the English meaning of the native-language word.
Gloss string `json:"gloss"`
// Definitions are the word's senses as written in the writer's own
// language — the monolingual half, for when the English gloss isn't enough.
Definitions []Meaning `json:"definitions,omitempty"`
// Phonetic is IPA for the native-language pronunciation; "" when absent.
Phonetic string `json:"phonetic,omitempty"`
}
// unknownDifficulty is the [Result.Difficulty] value meaning "no score",
// matching DreamDict's own -1 return.
const unknownDifficulty = -1
// GlossResult is the lightweight payload for the inline hover/select gloss: just
// the word and its Chinese translation, no definitions or synonyms. Kept small
// so the hover tooltip is instant and trivially cacheable.
type GlossResult struct {
Word string `json:"word"`
Gloss string `json:"gloss"`
// Reverse is the English meaning of the word read as one of the writer's
// own language — the tooltip's half of the both-directions rule (see
// [Reverse]). Empty unless the token is a word in her language too.
Reverse string `json:"reverse,omitempty"`
}
// maxSynonyms caps how many synonyms we hand the popover, even though the dataset
// stores up to ~50 per word — a long flat wall of words overwhelms more than it
// helps, especially for an ESL reader scanning for the right fit.
const maxSynonyms = 16
// maxDefinitions caps the senses shown so the popover stays a glance, not an
// essay.
const maxDefinitions = 4
// Lexicon holds the lazily-loaded datasets. The maps are populated once, on the
// first Lookup, behind a sync.Once so startup stays instant and a load error is
// remembered rather than retried on every request.
type Lexicon struct {
once sync.Once
loadErr error
defs map[string][][]string // word → [[pos, def, example], …]
synonyms map[string][]string // word → [synonym, …]
gloss map[string]string // word → Chinese gloss
phonetic map[string]string // word → IPA
}
// New returns a Lexicon. The datasets aren't read until the first Lookup.
func New() *Lexicon { return &Lexicon{} }
func (l *Lexicon) load() {
l.once.Do(func() {
if err := gunzipJSON(definitionsGz, &l.defs); err != nil {
l.loadErr = fmt.Errorf("load definitions: %w", err)
return
}
if err := gunzipJSON(synonymsGz, &l.synonyms); err != nil {
l.loadErr = fmt.Errorf("load synonyms: %w", err)
return
}
if err := gunzipJSON(glossGz, &l.gloss); err != nil {
l.loadErr = fmt.Errorf("load gloss: %w", err)
return
}
if err := gunzipJSON(phoneticGz, &l.phonetic); err != nil {
l.loadErr = fmt.Errorf("load phonetic: %w", err)
return
}
})
}
// Lookup returns the definition senses and synonyms for word. It first tries the
// word as written (lowercased), then a few simple morphological reductions
// (plurals, -ed/-ing/-ly) so "running" or "happily" still resolve. A word found
// in neither dataset yields a Result with empty lists (not an error).
func (l *Lexicon) Lookup(word string) (Result, error) {
l.load()
if l.loadErr != nil {
return Result{}, l.loadErr
}
norm := strings.ToLower(strings.TrimSpace(word))
res := Result{Word: word, Definitions: []Meaning{}, Synonyms: []string{}, Difficulty: unknownDifficulty}
if norm == "" {
return res, nil
}
res.Gloss = lookupGloss(l.gloss, norm)
// Phonetic is a word→string map like the gloss, so the same de-inflecting
// candidate walk resolves "running" → "run", etc.
res.Phonetic = lookupGloss(l.phonetic, norm)
if raw := lookupDefs(l.defs, norm); raw != nil {
for _, m := range raw {
res.Definitions = append(res.Definitions, toMeaning(m))
if len(res.Definitions) >= maxDefinitions {
break
}
}
}
if syns := lookupSyns(l.synonyms, norm); syns != nil {
if len(syns) > maxSynonyms {
syns = syns[:maxSynonyms]
}
res.Synonyms = append(res.Synonyms, syns...)
}
return res, nil
}
// Gloss returns just the Chinese translation for word (empty when absent). This
// is the fast path behind the inline hover/select gloss — it skips the
// definition and synonym datasets entirely.
func (l *Lexicon) Gloss(word string) (GlossResult, error) {
l.load()
if l.loadErr != nil {
return GlossResult{}, l.loadErr
}
norm := strings.ToLower(strings.TrimSpace(word))
return GlossResult{Word: word, Gloss: lookupGloss(l.gloss, norm)}, nil
}
// lookupGloss walks the candidate forms of a word and returns the first gloss
// hit (so "running"/"studies" resolve via the same de-inflection as defs/syns).
func lookupGloss(m map[string]string, word string) string {
if word == "" {
return ""
}
for _, c := range candidates(word) {
if v, ok := m[c]; ok {
return v
}
}
return ""
}
// lookupDefs / lookupSyns walk the candidate forms of a word and return the
// first dataset hit. They're separate (rather than a generic helper) only
// because the two maps have different value types.
func lookupDefs(m map[string][][]string, word string) [][]string {
for _, c := range candidates(word) {
if v, ok := m[c]; ok {
return v
}
}
return nil
}
func lookupSyns(m map[string][]string, word string) []string {
for _, c := range candidates(word) {
if v, ok := m[c]; ok {
return v
}
}
return nil
}
// candidates returns the lemma forms to try, in priority order: the word itself,
// then conservative de-inflections. This is deliberately lightweight — a full
// stemmer would over-reduce ("business" → "busy") and surface wrong entries; a
// handful of common English suffix rules covers the everyday cases without a
// dependency. Duplicates are fine (map lookup is cheap); order is what matters.
func candidates(word string) []string {
out := []string{word}
add := func(s string) {
if len(s) >= 2 && s != word {
out = append(out, s)
}
}
switch {
case strings.HasSuffix(word, "ies"): // studies → study
add(word[:len(word)-3] + "y")
case strings.HasSuffix(word, "es"): // boxes → box, wishes → wish
add(word[:len(word)-2])
add(word[:len(word)-1])
case strings.HasSuffix(word, "s"): // cats → cat
add(word[:len(word)-1])
}
if strings.HasSuffix(word, "ing") { // running → run, making → make
stem := word[:len(word)-3]
add(stem)
add(stem + "e")
add(undouble(stem))
}
if strings.HasSuffix(word, "ed") { // hoped → hope, stopped → stop
stem := word[:len(word)-2]
add(stem)
add(word[:len(word)-1])
add(undouble(stem))
}
if strings.HasSuffix(word, "ly") { // happily handled above via ies path? no — quickly → quick
add(word[:len(word)-2])
}
if strings.HasSuffix(word, "ily") { // happily → happy
add(word[:len(word)-3] + "y")
}
return out
}
// undouble collapses a doubled final consonant (stopp → stop, runn → run) so the
// -ed/-ing stems of doubled-consonant verbs resolve to their base form.
func undouble(stem string) string {
n := len(stem)
if n >= 2 && stem[n-1] == stem[n-2] {
return stem[:n-1]
}
return stem
}
// toMeaning maps a compact [pos, def, example] triple from the dataset onto the
// JSON-friendly Meaning. The dataset always stores three elements, but we guard
// the length so a malformed row can't panic.
func toMeaning(m []string) Meaning {
var out Meaning
if len(m) > 0 {
out.PartOfSpeech = m[0]
}
if len(m) > 1 {
out.Definition = m[1]
}
if len(m) > 2 {
out.Example = m[2]
}
return out
}
// gunzipJSON decompresses gz and decodes the JSON into v.
func gunzipJSON(gz []byte, v any) error {
r, err := gzip.NewReader(bytes.NewReader(gz))
if err != nil {
return err
}
defer r.Close()
data, err := io.ReadAll(r)
if err != nil {
return err
}
return json.Unmarshal(data, v)
}