Files
petal/scripts/build_ptpt_dictionary.py
T
prosolis ccb43e5a4d Phase 21: Petal learns to be an English+Portuguese pair
The plan said "Hunspell pt-PT vendored like en-US". Measuring that first is
what saved it: nspell expands affixes eagerly on construction, and European
Portuguese's 1,340 rules over 44,257 stems want over a gigabyte of browser
heap — ~340 MB for the first 12,000 entries, and no return at all after three
minutes on the whole file. So the expansion runs once at build time instead:
1,039,058 forms, 2.66 MB gzipped, read by the same nspell in 842 ms.

The obvious npm package would also have shipped the wrong language. Both
dictionary-pt and dictionary-pt-br carry VERO, the Brazilian word list, so
vendoring by name puts pt-BR spellings behind a pt-PT label — the drift
SUGGESTIONS §3 warns about, arriving through the packaging where no reviewer
can see it. The source is Projecto Natura's, and the build script now asserts
the fault lines (receção in, recepção out) before writing anything.

Spellcheck consults both dictionaries and flags only what both reject, which
is the no-detector answer to a pair with no script boundary. The word card
does the same in the other direction: "data" is a word in both languages, so
Petal shows both readings rather than guessing which she meant.

Writing the tests caught the one real bug — extendedAlphabet was a snapshot
while correct/suggest read live, and her dictionary arrives after English, so
every lookup would have resolved "cora" while the underlines were already
right.

Not done, and not claimed: the pack has not been read by a pt-PT speaker, and
the Piper voice is deferred with the deploy.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 12:43:02 -07:00

233 lines
9.5 KiB
Python

#!/usr/bin/env python3
"""Build Petal's browser pt-PT spelling dictionary from the LibreOffice Hunspell one.
Why this script exists at all
----------------------------
English is vendored the obvious way: `dictionary-en`'s `en.aff` + `en.dic` go
into web/public/dictionaries/en and nspell reads them in the browser. The plan
for Phase 21 said "Hunspell pt-PT vendored like en-US", and that turns out not to
work, for a measured reason.
nspell expands affixes **eagerly at load time** — it materialises every surface
form into a hash the moment you construct it. English gets away with this: ~50k
stems and a small rule set. European Portuguese does not. `pt_PT.aff` carries
1,340 affix rules (the full verb paradigm: six persons x a dozen tenses, plus
diminutives, plus productive prefixes) over 44,257 stems. Measured on this
machine, nspell needed ~340 MB of heap for the first 12,000 entries alone and had
not returned after three minutes on the whole file; extrapolated, it wants well
over a gigabyte. That is not something to hand a browser, still less a tablet.
So the expansion happens **here**, once, at build time, and the browser gets a
flat word list it can load with no affix machinery at all: 1,039,058 forms, 15 MB
of text, 2.7 MB gzipped, which nspell reads in ~0.8s using ~120 MB. The runtime
code path is then *identical* to English — same nspell, same interface — which is
the real prize. The aff we ship alongside keeps only the suggestion-shaping
directives (TRY/KEY/REP/MAP), so corrections still know that "cao" wants "ção"
and that a missing acute accent is a near miss.
Choosing the source
-------------------
npm's `dictionary-pt` is **not** European Portuguese. Both it and
`dictionary-pt-br` package VERO ("Verificador Ortográfico Livre", Brasil), so
vendoring the obvious npm name would have shipped Brazilian spellings under a
pt-PT label — exactly the pt-BR drift SUGGESTIONS.md §3 warns about, arriving
through the packaging rather than through the model.
The authentic dictionary is the Projecto Natura one (Universidade do Minho) that
LibreOffice ships and Debian packages as `hunspell-pt-pt`. Its aff declares
`LANG pt_PT`. Spot-checked against the built list, it accepts `receção`,
`húmido`, `telemóvel`, `autocarro`, `comboio`, `ótimo` and `pensámos`, and
rejects `recepção`, `úmido`, `ônibus` and `óptimo` — post-Acordo European
Portuguese, which is what a pt-PT writer should be held to.
Licensing: GPL-2 or LGPL-2.1 or MPL-1.1, (c) José João de Almeida, Rui Vilela,
Alberto Simões. The upstream copyright file is vendored beside the output.
Usage
-----
apt-get download hunspell-pt-pt # or take pt_PT.aff/.dic from LibreOffice
dpkg-deb -x hunspell-pt-pt_*.deb ptpt
python3 scripts/build_ptpt_dictionary.py \
ptpt/usr/share/hunspell/pt_PT.aff \
ptpt/usr/share/hunspell/pt_PT.dic \
web/public/dictionaries/pt-PT
Only the directives pt_PT.aff actually uses are implemented: PFX/SFX with
single-character flags, strip/append/condition, and cross-product. It carries no
compounding, no flag aliases and no NEEDAFFIX, so there is nothing else to
honour — the script asserts that rather than assuming it.
"""
import gzip
import os
import re
import sys
# Directives worth keeping in the shipped aff. These shape *suggestions*, not
# membership: TRY orders the alphabet the corrector tries, KEY knows which keys
# are adjacent, REP holds Portuguese-specific confusions (cao/ção, ss/ç), and MAP
# says an accented vowel and its bare form are the same letter for scoring —
# which is most of what an ESL writer gets wrong in Portuguese.
KEEP_DIRECTIVES = ("SET", "TRY", "KEY", "REP", "MAP", "WORDCHARS")
# Directives that would change which words are *accepted*. If a future upstream
# release starts using one, this script's output would silently disagree with
# the real dictionary, so it stops instead.
UNSUPPORTED = (
"COMPOUNDFLAG", "COMPOUNDMIN", "COMPOUNDRULE", "COMPOUNDBEGIN",
"ONLYINCOMPOUND", "NEEDAFFIX", "PSEUDOROOT", "CIRCUMFIX",
"FORBIDDENWORD", "AF", "AM", "FLAG",
)
def parse_aff(path):
"""Return (prefix rules, suffix rules, cross-product flags) keyed by flag."""
with open(path, encoding="utf-8") as fh:
lines = fh.read().splitlines()
for line in lines:
head = line.split()[0] if line.split() else ""
if head in UNSUPPORTED:
raise SystemExit(
f"{path}: unsupported directive {head!r} — this expander only "
"handles plain PFX/SFX affixation, and honouring it would change "
"which words are accepted. Extend the script before shipping."
)
pfx, sfx, cross = {}, {}, {}
i = 0
while i < len(lines):
parts = lines[i].split()
if parts and parts[0] in ("PFX", "SFX"):
kind, flag, cross_flag, count = parts[0], parts[1], parts[2], int(parts[3])
table = pfx if kind == "PFX" else sfx
cross[flag] = cross_flag == "Y"
rules = table.setdefault(flag, [])
for j in range(1, count + 1):
p = lines[i + j].split()
strip = "" if p[2] == "0" else p[2]
# The appended text may carry its own continuation flags after a
# slash (append/FLAGS). We drop them: honouring them would mean
# affixing an affixed form, which pt_PT.aff does not do.
append = "" if p[3] == "0" else p[3].split("/")[0]
cond = p[4] if len(p) > 4 else "."
anchored = ("^" + cond) if kind == "PFX" else (cond + "$")
rules.append((strip, append, re.compile(anchored)))
i += count + 1
continue
i += 1
return pfx, sfx, cross
def apply_suffix(word, rules):
out = []
for strip, append, cond in rules:
if strip and not word.endswith(strip):
continue
if not cond.search(word):
continue
stem = word[: len(word) - len(strip)] if strip else word
out.append(stem + append)
return out
def apply_prefix(word, rules):
out = []
for strip, append, cond in rules:
if strip and not word.startswith(strip):
continue
if not cond.search(word):
continue
stem = word[len(strip):] if strip else word
out.append(append + stem)
return out
def expand(aff_path, dic_path):
pfx, sfx, cross = parse_aff(aff_path)
forms = set()
with open(dic_path, encoding="utf-8") as fh:
fh.readline() # leading entry count, not a word
for raw in fh:
entry = raw.strip().split("\t")[0] # drop morphological fields
if not entry:
continue
word, _, flagstr = entry.partition("/")
word = word.strip()
if not word:
continue
flags = set(flagstr.strip())
forms.add(word)
for f in flags:
if f in sfx:
forms.update(apply_suffix(word, sfx[f]))
prefixed = []
for f in flags:
if f in pfx:
prefixed.extend(apply_prefix(word, pfx[f]))
forms.update(prefixed)
# Cross-product: a prefixed form may also take a suffix, but only
# when both rules are declared cross-product ("Y").
for f in flags:
if f in pfx and cross.get(f):
for base in apply_prefix(word, pfx[f]):
for g in flags:
if g in sfx and cross.get(g):
forms.update(apply_suffix(base, sfx[g]))
return forms
def shipped_aff(aff_path):
keep = []
for line in open(aff_path, encoding="utf-8").read().splitlines():
head = line.split()[0] if line.split() else ""
if head in KEEP_DIRECTIVES:
keep.append(line)
return "\n".join(keep) + "\n"
# Words the built list must accept, and must reject, before it is written. These
# are the pt-PT/pt-BR fault lines: post-Acordo spellings, the European lexicon,
# and the first-person-plural preterite accent that only pt-PT writes. A source
# dictionary that fails these is not the one this script is for.
MUST_ACCEPT = ("receção", "húmido", "telemóvel", "autocarro", "comboio",
"ótimo", "pensámos", "escrevêssemos", "jardim")
MUST_REJECT = ("recepção", "úmido", "ônibus", "óptimo")
def main(aff_path, dic_path, out_dir):
forms = expand(aff_path, dic_path)
missing = [w for w in MUST_ACCEPT if w not in forms]
present = [w for w in MUST_REJECT if w in forms]
if missing or present:
raise SystemExit(
"this does not look like European Portuguese: "
f"missing {missing}, unexpectedly present {present}"
)
os.makedirs(out_dir, exist_ok=True)
ordered = sorted(forms)
body = f"{len(ordered)}\n" + "\n".join(ordered) + "\n"
dic_out = os.path.join(out_dir, "pt-PT.dic.gz")
# mtime=0 so rebuilding identical input produces an identical file — a
# vendored asset that changes on every build is noise in the diff.
with gzip.GzipFile(dic_out, "wb", compresslevel=9, mtime=0) as fh:
fh.write(body.encode("utf-8"))
aff_out = os.path.join(out_dir, "pt-PT.aff")
with open(aff_out, "w", encoding="utf-8") as fh:
fh.write(shipped_aff(aff_path))
print(f"{len(ordered)} forms -> {dic_out} "
f"({os.path.getsize(dic_out) / 1e6:.2f} MB gzipped)", file=sys.stderr)
if __name__ == "__main__":
if len(sys.argv) != 4:
raise SystemExit(__doc__.strip().splitlines()[-1])
main(*sys.argv[1:4])