`pair_lang` had always been answering a second question nobody asked: it says which two languages, and every surface built on it assumed English was the one being learned. That is why hanzi is never tokenized, never spell-checked, never glossed — correct for a Mandarin native practising English, backwards for an English native practising Mandarin. `users.direction` (migration 0016) separates the two questions; a `zh-learner` pair code would have been cheaper and would have made two directions of one pair look like two unrelated languages to every query. Segmentation is what replaces `wordAt` where there are no spaces: a shortest-path walk over log-probabilities, 232 ms and 14 MB for 188,522 words. The browser gets the word list because segmentation runs on hover; the server keeps the whole dictionary. Their coverage gates come out opposite on purpose — the client list is frequency-gated because the segmentation is measurably identical without the tail, and the dictionary is gated by nothing, because its only power is to explain and the word a learner stops on is the rare one. The 错别字 pack is 24 confusable pairs behind two mechanical gates. One admits a pair only if the wrong form is not a dictionary word and the right form is, which is why it refuses 自已 for 自己 — a real error whose wrong form is a headword. The other asks the segmenter whether the two characters already belong to two different words, without which 自己经常, 睡觉的时候 and 不知到底 would all be corrupted silently into text still made of real characters. Not deployed (this carries a migration), not seen in a browser, and no account has ever been in the learner direction. The IME composition guards were in scope and are not done — see BUILD_PLAN Phase 26.
122 lines
5.2 KiB
Go
122 lines
5.2 KiB
Go
package lexicon
|
|
|
|
// A word lookup used to mean exactly one thing: the embedded datasets, which
|
|
// speak English and Mandarin and nothing else. That was fine while Petal had
|
|
// one writer. It stops being fine the moment a pt-PT writer right-clicks a
|
|
// word and gets a Chinese gloss.
|
|
//
|
|
// So the lookup becomes a seam. A [Provider] answers the same two questions the
|
|
// popover and the hover tooltip have always asked; which provider answers them
|
|
// depends on the writer's language pair, and [Set.For] is the only place that
|
|
// decision is made.
|
|
|
|
// Provider answers word lookups for one writer. The embedded datasets and
|
|
// DreamDict both satisfy it, and both treat a word they don't carry as an empty
|
|
// result rather than an error — a miss is an ordinary outcome of looking a word
|
|
// up, not a failure.
|
|
type Provider interface {
|
|
// Lookup returns the full popover payload: gloss, phonetic, definitions,
|
|
// synonyms, and whatever extras the provider carries.
|
|
Lookup(word string) (Result, error)
|
|
// Gloss returns just the writer's-language translation. It is the hover
|
|
// tooltip's fast path and skips everything else.
|
|
Gloss(word string) (GlossResult, error)
|
|
}
|
|
|
|
// LangZh is the one pair language still served by the embedded datasets. Every
|
|
// other pair goes to DreamDict — see [Set.For] for why zh is held back.
|
|
const LangZh = "zh"
|
|
|
|
// langEN is the language DreamDict is asked about for definitions, synonyms and
|
|
// pronunciation. English is always the *target* language of the pair — what
|
|
// varies is the language the gloss is written in.
|
|
const langEN = "en"
|
|
|
|
// Set holds every provider Petal can serve a lookup from and picks between them
|
|
// by pair language. One Set is shared by the whole process: the embedded
|
|
// datasets load once, and dict.db is one read-only handle.
|
|
type Set struct {
|
|
embedded *Lexicon
|
|
// dream is nil when dict.db was not deployed. That is a supported state,
|
|
// not an error — see [Set.For].
|
|
dream *DreamDict
|
|
}
|
|
|
|
// NewSet returns a Set backed by the embedded datasets and, when dream is
|
|
// non-nil, DreamDict. Passing a nil dream is how Petal runs without dict.db.
|
|
func NewSet(dream *DreamDict) *Set {
|
|
return &Set{embedded: New(), dream: dream}
|
|
}
|
|
|
|
// HasDreamDict reports whether a dict.db is open. Only startup logging and
|
|
// tests care; a handler never asks, because [Set.For] always returns something
|
|
// usable.
|
|
func (s *Set) HasDreamDict() bool { return s.dream != nil }
|
|
|
|
// Contents describes what the open dict.db actually holds, for the startup log.
|
|
// With no dictionary it says so rather than returning an empty string, because
|
|
// a blank in a log line is indistinguishable from a bug in the log line.
|
|
func (s *Set) Contents() string {
|
|
if s.dream == nil {
|
|
return "no dict.db — embedded datasets only"
|
|
}
|
|
return s.dream.Contents()
|
|
}
|
|
|
|
// For returns the provider that should answer lookups for a writer whose pair
|
|
// language is lang.
|
|
//
|
|
// Three rules, in order:
|
|
//
|
|
// zh — and an empty code, which is what a pre-Phase-16 row reads as — stays on
|
|
// the embedded ECDICT gloss. Not because DreamDict lacks Chinese (it has
|
|
// CC-CEDICT), but because that path is in daily use by a real writer and the
|
|
// two have not yet been compared on her actual lookups. Switching it is a
|
|
// quality decision, and it hasn't been made.
|
|
//
|
|
// Any other pair goes to DreamDict, which is the only source that has pt-PT,
|
|
// French or Spanish at all.
|
|
//
|
|
// If dict.db was never deployed, a non-zh writer falls back to the embedded
|
|
// datasets with the gloss suppressed. This is the interesting case: the naive
|
|
// "no data" answer would blank the popover entirely, when in fact the English
|
|
// half of it — definitions, synonyms, phonetic — is compiled into the binary
|
|
// and perfectly correct for her. Only the translation is missing, so only the
|
|
// translation goes missing. A failed dictionary deploy costs her the gloss, not
|
|
// the dictionary.
|
|
func (s *Set) For(lang string) Provider {
|
|
if lang == "" || lang == LangZh {
|
|
return s.embedded
|
|
}
|
|
if s.dream != nil {
|
|
return dreamProvider{dict: s.dream, native: lang}
|
|
}
|
|
return glossless{s.embedded}
|
|
}
|
|
|
|
// glossless serves the embedded datasets with the Chinese gloss stripped, for a
|
|
// writer who does not read Chinese. Handing her the zh gloss would be worse
|
|
// than handing her nothing: an empty field reads as "not found", where the
|
|
// wrong language reads as Petal being broken.
|
|
type glossless struct{ inner Provider }
|
|
|
|
func (g glossless) Lookup(word string) (Result, error) {
|
|
res, err := g.inner.Lookup(word)
|
|
res.Gloss = ""
|
|
return res, err
|
|
}
|
|
|
|
func (g glossless) Gloss(word string) (GlossResult, error) {
|
|
return GlossResult{Word: word}, nil
|
|
}
|
|
|
|
// Hanzi answers a Chinese-word lookup from the embedded CC-CEDICT map.
|
|
//
|
|
// It is on the Set rather than on [Provider] because it is not the same
|
|
// question the other two ask. Lookup and Gloss vary by pair — which is why they
|
|
// are behind an interface with two implementations — while this one is asked of
|
|
// Chinese or not at all: the learner direction exists for exactly one pair (see
|
|
// auth.learnerPairs), and DreamDict's own CC-CEDICT would be a second copy of
|
|
// the same dictionary, chosen by a rule with one branch.
|
|
func (s *Set) Hanzi(word string) (HanziResult, error) { return s.embedded.Hanzi(word) }
|