Compare commits

...
Author SHA1 Message Date
prosolis 8f2ad34a10 Four ways the two scripts weren't the same app, and a smaller cat
A review of the pair work found the seams — every one of them a place where
the Chinese half was written and the older Latin half was left standing.

The right-click menu still asked the Latin tokenizer whether there was a word
under the pointer, so right-clicking a hanzi opened the browser's own menu
instead of the card. Hover, long-press and Ctrl+D had all moved to the shared
resolver; this one hadn't, and it is the surface the segmenter's own header
names first.

isHan is a property escape precisely so the extension blocks are covered, and
then every call site handed it one UTF-16 code unit — half a surrogate pair
for anything above the BMP, which \p{Script=Han} rightly says is not Han. The
run split in two around the character and the words either side stopped being
looked up. The walk, the scan and wordAt now step by code point, the regex is
anchored, and the test that passed by accident (unanchored, so it searched a
two-unit string rather than testing one character) is joined by one that
would have failed.

The pair picker sent the pair alone. The server validates pair and direction
as one decision and refuses a learner direction for a pair it has no word
list for — so an English speaker learning Chinese could not move to French at
all: every button failed with the generic message. It now names both, keeps
her direction where the target pack has a learner side, and returns her to
learning_en where it does not. Routed through useSession rather than the
picker's own api call, so me.direction — which decides whether the word list
stays loaded — moves with it.

UpdateMe answered every error from Get with 401. A SQLite fault on a PATCH
would have tripped the client's session interceptor and thrown a writer into
the signed-out overlay while her session was fine. Only a missing row means
not signed in, which is the distinction SetPair already made below it.

emitCommittedRef was assigned during render and called later from
compositionend; a render React discards must not leave its closure behind for
a DOM event.

And the kitten is 10% smaller — one clamp, three terms, everything else
calc()s off it.

vitest 297/297, tsc, go build/vet/test clean.
2026-07-28 20:03:45 -07:00
prosolis 5659312358 A real IME, a real candidate window, and 11.6 seconds of holding still 2026-07-28 19:39:21 -07:00
prosolis a216614c81 Phase 27 seen in a browser: held, released, and the save never held with it 2026-07-28 19:32:14 -07:00
prosolis c348a9b8ae The keystroke that isn't one: IME composition guards
Phase 26 scoped these and left them unbuilt, naming them as the likeliest
thing to be wrong the first time anyone types Chinese into Petal for real.

A composition is not a keystroke: the pinyin goes into the document as it
is typed, a candidate window sits over it, and all three decoration layers
recompute from the live document on every change — rewriting the DOM around
the node the browser is composing in, which is what eats half-typed input.

The layers now hold their redraws rather than skip them: a rebuild that falls
due mid-composition marks itself stale and its decorations are mapped through
the transaction, so they travel with the text and land correct the moment the
composition ends. The flag is read from the state before the transaction, so
the answer doesn't depend on plugin ordering; the end transaction is the one
deliberate exception, or nothing would ever release. The release is a
macrotask late because a custom handleDOMEvents handler runs before
ProseMirror's own and ProseMirror flushes the composition's last changes in a
microtask — so the held rebuild sees the committed hanzi, not the pinyin it
replaced.

Input rules needed no guard (Tiptap already returns early while composing),
which was checked rather than assumed: pinyin uses an apostrophe as a
syllable separator and Typography rewrites every ' into a curly one.

The save is deliberately not gated and the analysis is. A tablet keyboard can
hold one composition open for a whole sentence, and Petal never makes writing
wait for anything — so EditorChange carries the flag, auto-save ignores it,
and the checkpoint, rule pack and companion wait for the word to commit. One
more change is emitted the instant it does, so nothing is skipped.

Four places were taking keys that belong to the IME: the Find bar, the tag
picker, Ask Petal's chat box, and distraction-free mode's global Escape.

vitest 296/296, tsc, vite, go build/vet/test clean. Not verified with a real
IME — no browser or IME here, and that is the half the tests cannot reach.
2026-07-28 19:20:11 -07:00
prosolis 77f284f65c The zh pair's other direction, and a rule pack that mostly says no
`pair_lang` had always been answering a second question nobody asked: it
says which two languages, and every surface built on it assumed English
was the one being learned. That is why hanzi is never tokenized, never
spell-checked, never glossed — correct for a Mandarin native practising
English, backwards for an English native practising Mandarin.
`users.direction` (migration 0016) separates the two questions; a
`zh-learner` pair code would have been cheaper and would have made two
directions of one pair look like two unrelated languages to every query.

Segmentation is what replaces `wordAt` where there are no spaces: a
shortest-path walk over log-probabilities, 232 ms and 14 MB for 188,522
words. The browser gets the word list because segmentation runs on hover;
the server keeps the whole dictionary. Their coverage gates come out
opposite on purpose — the client list is frequency-gated because the
segmentation is measurably identical without the tail, and the dictionary
is gated by nothing, because its only power is to explain and the word a
learner stops on is the rare one.

The 错别字 pack is 24 confusable pairs behind two mechanical gates. One
admits a pair only if the wrong form is not a dictionary word and the
right form is, which is why it refuses 自已 for 自己 — a real error whose
wrong form is a headword. The other asks the segmenter whether the two
characters already belong to two different words, without which 自己经常,
睡觉的时候 and 不知到底 would all be corrupted silently into text still
made of real characters.

Not deployed (this carries a migration), not seen in a browser, and no
account has ever been in the learner direction. The IME composition
guards were in scope and are not done — see BUILD_PLAN Phase 26.
2026-07-28 19:04:53 -07:00
prosolis 9224c44fff The es pair, and a dictionary that was quietly Spain's
Phase 25. Spanish was never built — the groundwork was all [x] (DreamDict
data, the prompt language, the L1 rule gating, TTS env-discovery), which is
why the plan read as though it had shipped. shippedPairs was the honest
answer: the server had been refusing es on purpose.

The langpack is neutral Latin American, chosen with the user: tú, ustedes,
no vosotros, and the pan-American half of every vocabulary split. A vitest
greps for the peninsular twins the way fr is greped for québécismes —
including coger, which is not merely regional but obscene through most of
Latin America.

The dictionary is the story. Debian's hunspell-es symlinks twenty country
codes to one file, which reads as pan-Hispanic; RLA publishes twenty-four
builds per release, one per country plus a generic es that is the union,
and Debian ships peninsular es_ES. The 58,622-form gap is essentially
voseo, so the first version of this commit underlined vení and tenés as
misspellings and called it a considered gap.

The MUST_ACCEPT list was written to catch exactly that and structurally
could not: it asserted the pan-Hispanic vocabulary, and every RLA variant
carries the full pan-Hispanic vocabulary — only the paradigms are
localised. The REP table cited as the second witness is shared by all
builds too. Two independent-looking proofs, neither able to distinguish
anything, agreeing with each other.

The profile now demands what discriminates, each verified against the build
it targets: voseo rejects es_ES and Debian, vosotros rejects es_MX, and
arepa/chévere/bacán reject es_AR, which has both paradigms and would
otherwise pass. 717,640 forms, 1.74 MB gzipped, 762 ms / 97 MB in a real
nspell. fr and pt-PT rebuild byte-identical from their own upstream debs,
so the shared script still means what it meant.

Shipping the union is fr's call arrived at from the other side: coût and
cout are both correct French, tienes and tenés are both correct Spanish.
The dictionary holds every variety because underlining is all it can do;
the copy picks a register because speaking requires one.

Reviewed by four models at the usual >=2-of-4 threshold, 5 of 27 findings
applied — one catching the bedtime proverb as fr's Qui dort dîne calqued
into Spanish, gloss and all, which is the rule the fr header states. One
below-threshold finding (a missing ¡, seen by 1 of 4 because an absent
opening mark has no closing ! to look wrong against) was applied and turned
into an assertion instead: the suite now rejects any native line that
closes ? or ! without opening one.

piper-es on es_MX-ald-medium, not the es_ES-davefx-medium the plan named —
six of Piper's nine Spanish voices are peninsular, so the obvious pick was
the pt-PT trap through a different door.

go build/vet/test, tsc, vite, vitest 251/251.

Not deployed, not seen in a browser, not read by a native speaker, and no
es account exists.
2026-07-28 18:25:59 -07:00
prosolis 39d4e4770a Merge feat/accept-all-category: a whole category in one click, and one undo 2026-07-28 17:30:51 -07:00
prosolis bd92cdc9b6 A whole category accepted in one click, and one undo
Five article fixes were five clicks, five confetti bursts and five undo
steps. "Accept all Tidy-up (5)" makes them one of each.

The single undo decided the implementation: every replacement goes into one
Tiptap chain, which applies as one transaction and so undoes as one history
event. That only works if the spans can't move under each other, so the
plan resolves every span against the document as it stands and applies them
last-first.

Three outcomes rather than one, because a batch that quietly dropped a card
would be reporting edits it never made: a span she already fixed herself is
settled without an edit (what a single Accept does too), and a card quoting
the same words as one already taken is left on screen, since findRange
would resolve both to the same place.

The control sits on the first card of its kind — the rail can't carry a
category header, its cards are anchored to their own sentences — and only
when the category has company. It is outlined rather than filled: it acts
on cards she can't see from where she's standing.
2026-07-28 17:30:44 -07:00
prosolis 1acc23244e UX review: item 8's two are done, and this stack is deliberately not deployed 2026-07-28 17:07:46 -07:00
prosolis 40de65b3d1 Merge feat/settled-spans: a dismissed card stays dismissed, and the bar counts the petals 2026-07-28 07:27:55 -07:00
prosolis b2d50e9136 A dismissed card stays dismissed, even offline
The server already suppressed every span she had accepted or dismissed, on
both the LLM reconcile and the mechanics pass, with tests either side. What
had no memory was the half that never asks it: item 3b's rule pack renders
250 ms after a keystroke with no network, and its record of "she already
answered this" was a set cleared on every document switch and added to only
for cards dismissed while still provisional.

So dismissing a persisted rule-pack card recorded nothing client-side and the
next keystroke put it straight back until the server's reply removed it again;
and after a reload the client knew nothing at all — permanently so with the
server unreachable, which is the case the rule pack exists for.

GET /docs/{id}/settled hands over the normalized originals of the document's
actioned rows, scoped through documents because an original quotes her
sentence. The client seeds a SettledSpans from it on open and adds to it for
every card that leaves, keyed on the original alone the way the server keys
it. The load adds rather than assigns, so a dismissal made while it is in
flight survives it.

normalizeForDedup now exists in both languages, compared across a network
boundary, so the same nine cases are asserted on both sides and each test
names the other.

Also: the status-bar count — "🌸 5片花瓣待打磨 · 5 petals to polish" beside the
word count, from the packs, hidden at zero. An empty rail already says nothing
is waiting; a badge announcing it after every check is a verdict, which the
review's non-goals rule out.

Verified in Chrome at 1517x810 with the server killed: a new violation was
detected, underlined and counted with no network, while the dismissed span
stayed gone.
2026-07-28 07:27:49 -07:00
prosolis 1aa3a14030 Merge fix/rail-follows-mode: the rail follows the mode, and Ask Petal answers in both languages
Item 7 (the rail is a mode, not a screen size; click opens the anchored
card even with the rail up) and item 6 (bilingual Ask Petal answers with
room to read) from the 2026-07-27 UX review.
2026-07-28 07:04:23 -07:00
prosolis 978cb80642 Ask Petal answers in both languages, with room to read
The tutor prompt said "never mix languages in a single response" and
mirrored the language of the question, so asking in English — which she
does, because she is practising — returned the one explanation surface
that gives nothing in her own language. It now answers in both, pair
language first, halves separated by a blank line.

Which half is the safety net and which is the lesson depends on who is
writing: the pair is (English + X) and Petal is used from both ends, so
the prompt asks for both and says it doesn't know which way round.

The split is a rendering nicety, never a parse the reply depends on: a
half-streamed reply is all one half, a model that ignores the
instruction renders as one block, and nothing is ever dropped.

For the height, the first attempt clamped the box to the room left below
the anchored card so it could never overhang — measured, that gave 176px
against a 442px answer, worse than the 220px it replaced. The card's own
chrome spends ~290px of an 810px window, so "fits below the word" and
"room to read" are not both available. The ceiling is now a flat 50vh and
the overhang is made navigable instead, per item 4: the card reports its
reach like the rail already does, the column grows, and the page can
scroll to the actions below it.
2026-07-28 07:04:16 -07:00
prosolis f082a930cb The rail is a mode, not a screen size
Item 7 said to confirm before building, and confirming is what mattered.

The rail's 348px threshold is measured against a fixed 720px column centred
in the pane. The doc-list sidebar is 280px, so at her 1517px viewport the
right margin is 258 with it open and 406 without — either side of the
threshold. What moves between them is distraction-free mode, which engages on
its own when the editor takes focus. The rail therefore appears when she
starts writing and disappears when she stops; items 4 and 5 disagreed about
whether it exists at 1517px only because they caught it in different states.

Re-centring a fixed-width column changes its position and not its size, so
the wrapper's ResizeObserver reported nothing and no window resize fired.
railEnabled kept whatever value it last had. Leaving distraction-free with
the rail up left a 300px column in a 266px margin: overhanging the viewport
by 66px, cards clipped mid-sentence, the page scrolling sideways. Entering it
with the rail down opened 406px of margin and put nothing in it. Both
persisted until something else happened to resize the window.

Observe the scrollport too — it spans the pane, so it resizes whenever the
chrome around the editor does. That covers any future chrome that moves the
editor, which threading focusMode down as a prop would not.

Clicking a highlight now opens the anchored card even when the rail is up.
That is the item's own acceptance criterion and was previously false by
design; the measured distance from the first flagged span to its rail card is
651px, not the ~400 the review guessed. Hover still defers to the rail, since
the reasoning against an unbidden second card was about hover and still
holds — but a click is her asking to deal with that word. The rail card glows
instead of expanding, so nothing is ever open twice.

Verified in Chrome at the review's own 1517x810, driving the rule pack from
item 3b so no model was involved: the rail follows the mode in both
directions with no resize event anywhere; the popover lands 6px under the
word with the full explanation, Ask Petal, Accept and Dismiss; accepting from
it applied the edit and took the rail 6 cards to 5, leaving the rest with
their ids, positions and wording intact.

railFit.test.ts pins the threshold to the margins actually measured. The
observer wiring has no unit test and can't have a useful one: jsdom has no
layout, so every rect is zero and the rail branch is unreachable there. That
half is browser-verified only, and the doc says so.
2026-07-28 06:34:27 -07:00
prosolis ce5de68b9b UX review: item 5 is live, and what the backup nearly missed
Stopping the container does not checkpoint the WAL: petal.db on disk was four
hours stale while 2 MB of her writing sat in petal.db-wal. A lone
`cp data/petal.db` would have backed up the wrong day — and 0015 is the first
migration here that rebuilds a table, so that backup was the one that mattered.
Recorded for the next one.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-28 00:35:45 -07:00
prosolis b23c5a9a13 Merge feat/translate-card: her own language gets its own card
When she reaches for Chinese mid-sentence, the card now says 翻译 · Translate
rather than Clarity. Petal already found the span and already rendered it into
English; the type was the whole gap, and it is derived from the span rather
than asked of the model so it can't drift between passes.

Carries migration 0015 (suggestions table rebuild for the extended type CHECK)
and two findings from the running page: the inline underline needs a per-type
colour rule or it renders invisibly, and at her viewport the rail is disabled —
the inline hover card is what she sees.

Green: go vet, go test ./..., tsc --noEmit, 199 vitest tests.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-28 00:32:04 -07:00
prosolis 25e415daa2 When she writes in Chinese, say Translate — not Clarity
She reaches for her own language mid-sentence when English won't come, and
Petal already handled it: it found the span and rendered it into English. It
just filed the result as a Clarity fix, so the pair model's flagship moment
read as tidying up her Chinese.

The type is now derived from the span rather than asked of the model. A type is
structural, and a model that re-reasons every pass would drift between labels
for a sentence nobody had touched — the instability the last session spent
itself removing. The label the model volunteers is still ignored.

Only the grammar checkpoint can be promoted. A pass with a forced type owns its
family: voice reads paragraphs for tone and its rows carry no replacement, so a
"translation" there would be a card offering nothing to accept.

zh is a different script and counting Han runes is close to certain. The Latin
pairs share an alphabet with English and get none of that, so they fall back to
function words and need two before Petal claims anything — with every word that
is also English left out, even the common ones. The heuristic is justified by
how cheap being wrong is: it changes a coloured pill, and nothing else.

The pill is the one bilingual type name in the rail. Every other type stays
English because those are the terms she is learning; this card's whole subject
is her own language. And it stops truncating its two lines — elsewhere the diff
is a word and the explanation is what she reads, but here the two sentences are
the card.

Two things only the running page could report. The inline underline was
invisible: the decoration carries a per-type class and the base rule is a
transparent border, so a type with no colour rule gets no mark at all. And at
1517×810 with the document list open there is no rail — the margin is 258 where
railEnabled wants 348 — so what she gets is the inline hover card. Item 7 is
written the other way round.

Migration 0015 rebuilds the suggestions table for the CHECK, which makes it the
first one here that could quietly drop her rows; there is a test that carries
every column, both timestamps and both indexes across it.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-28 00:09:42 -07:00
prosolis 3bcc967f51 UX review: the stack is deployed, and the handoff says so
Three sessions running, the handoff opened with "nothing deployed" and
closed by recommending item 3b — which had shipped two sessions earlier.
Both are now false, and a stale recommendation is worse than none: it sends
the next session to re-scope finished work.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 23:40:20 -07:00
prosolis 963fc1754d Merge the suggestion-loop stack: instant rules, stable cards, reachable rail
Three sessions of UX_REVIEW work land together, because they are one
change to how suggestions arrive and sit:

- item 3b — the deterministic rule pack renders on its own 250 ms fuse
  instead of waiting behind the LLM's 4 s checkpoint and a round-trip.
- item 2 — passes reconcile instead of replacing, so an untouched card
  keeps its id, its arrival chime and its original explanation, and an
  unchanged document doesn't call the model at all.
- item 4 — the rail's overhang becomes real scrollable page, with the
  prose pinned bottom-anchored so the sentences the lower cards flag
  stay on screen.

Green: go test ./... , tsc --noEmit, 195 vitest tests.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 23:37:56 -07:00
prosolis de251ceae2 Make the suggestion rail's overhang reachable, and keep the text in view
The margin rail hangs off an absolutely-positioned column, so its cards add
no layout height. On her live document that meant four 173px cards anchored
inside 126px of text: a 714px stack over a page whose scrollHeight equalled
its clientHeight. The lower cards weren't far from their sentence, they were
off-screen with nothing to scroll.

The rail now reports how far its resolved stack reaches and the wrapper takes
that as a minimum height, so the space those cards occupy is real, scrollable
page. minHeight never shrinks the column, so a rail that fits beside its text
is unaffected.

Scrolling into that space would have carried every sentence off the top, so
the prose is pinned while the stack overhangs it. The offset is
min(0, port - content): prose shorter than the viewport pins at the top,
taller prose pins by its bottom edge, keeping the last lines visible — those
are the ones the overhanging cards flag.

The prose box has to stay at its natural height. Keeping the old h-full made
it measure the wrapper this change had just grown, reporting the cards'
height as the text's own, so the pin could never trip.

Verified in a browser at the review's 1517x810, driven offline by the rule
pack: 8 cards over 95px of prose gained 675px of scroll where there was none,
the last card lands fully in view with the text still on screen, tall prose
pins bottom-anchored without disturbing ordinary scrolling, and hover-linking
still glows the right span.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 23:07:09 -07:00
prosolis 10e8aef86c Stop regenerating the world on every check
A card vanishing and coming back seconds later, with different words, was
never about latency: every pass deleted its whole family and re-inserted
it, so each round minted new row ids. The rail keys on suggestion.id, so a
full remount was guaranteed — new id, new created_at (hence the re-fired
chime), and a fresh explanation from a model that re-reasons every time it
is asked. One unchanged mistake carried three different explanations in a
single sitting.

Passes now reconcile instead of replace. A re-proposed edit keeps its row:
its id, its created_at, and the wording she has already read. And the
grammar checkpoint stops asking about sentences nobody touched — the
document is split into hashed sentences, checked_chunks records which ones
a family has read, and only the difference is sent. When nothing changed
it doesn't call the model at all, and doesn't spend its rate-limit slot on
having done nothing.

The tone is part of a sentence's identity: cached advice was written for
the old register, so switching doc type re-reads every line.

replaceMechanics reconciles too, which mattered more than expected — the
rule pack fires 250 ms after a keystroke, so it was re-minting every local
card's id several times a sentence.

Only the grammar checkpoint is chunked. Voice is a property of the whole
document, and the collocation coach is a button she pressed asking for a
fresh read.

No client change was needed; stable ids were the whole of it.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 22:46:12 -07:00
prosolis c33de1175b Render deterministic rule hits instantly, not on the LLM's clock
The rule pack in prose.ts already found "a apple" — articles,
pluralAfterNumber, subjectVerbAgreement, uncountables are all there, and
they already surface as real mechanics cards. But mechanicsFindings only
ran inside runCheck, behind the same 4s checkpoint debounce as the model,
and only reached the screen via the server's reply. A free, instant,
offline-capable detection was being delivered on an LLM-shaped delay.

The rule pack now runs on its own 250ms fuse and renders its findings with
no network at all, as provisional cards. The mechanics submit follows; its
reply is authoritative and clears them. If the reply never comes — offline,
server down — the cards simply stay, which is the whole point of having
rules that need no model.

Provisional cards are keyed by wording rather than position, so one can't
flicker into a duplicate of its own persisted twin while she types around
it. resolveServerId maps a card to the row the API can act on, awaiting the
in-flight submit, so accepting inside that window still records the keep and
plants its word in the garden instead of being quietly dropped; null means
there is no row and the edit has landed regardless. Findings she actions
while provisional are remembered client-side, because the detector has no
memory between runs. runCheck no longer re-submits what the fast pass
already filed — it's the catch-up path for when that submit failed.

The arrival chime keys rule-pack cards by wording too, so a finding doesn't
chime once as provisional and again as persisted.

Not done, deliberately: no distinct style for unconfirmed local hits. The
rail renders both engines identically on purpose, and a provisional card now
lives for one LAN round-trip.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 22:28:37 -07:00
prosolis ba06d904f0 UX review: correct the item 3b handoff advice
The previous handoff sent the next session off to read
feat/mechanics-deterministic-pass and feat/calm-suggestions as unmerged
branches. Both are in main and have been for a while — that came from
misreading `git branch -vv` tracking info as merge status.

It matters because it inverts the advice. The deterministic rules engine
(prose.ts) already ships, already emits exact-span fixes as suggestion
cards under a 'mechanics' family, and already suppresses re-edits of
settled sentences. So item 3b's remaining work is most likely the
latency/ordering half — render local hits before the LLM pass — not
writing a rules engine. Point at the code and the two commits instead.

Same for item 5: internal/suggestions/translate.go already exists.

Also record that every topic branch was fully merged and has now been
deleted locally and on origin; main is the only branch left.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 22:20:54 -07:00
prosolis ea14eb5e88 UX review: handoff notes for the next session
Records what shipped and is live, what was closed without code (items 1
and 5), what's untouched, and the suggested next step — item 3b, with a
warning to read grammarLite.test.ts and the unmerged
feat/mechanics-deterministic-pass and feat/calm-suggestions branches
before writing a new rules engine.

Also writes down how to instrument a production build, since item 1
looked airtight in source and was wrong: fiber-walk from .ProseMirror to
the Tiptap editor, then read the prosemirror-history state directly.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 22:17:26 -07:00
prosolis aac15b5ac5 Merge fix/companion-yields-to-cards: the kitten gets out of the way
Fades, shrinks and goes click-through when suggestion cards or the
History/Garden drawers reach its corner, so nothing it sits on top of is
ever unreadable or unclickable.

Also lands the 2026-07-27 UX review doc, with items 1 and 5 corrected
against the live build.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 22:12:47 -07:00
prosolis be9aa13287 Kitten yields to panels too; redo bug not reproducible
The overlap hook only watched .petal-rail-card, so the History and Garden
drawers still sat under the mascot — with a real control ("写作证明 ·
Writing passport") buried under the halo on the live build.

Match [role="dialog"][aria-modal="true"] as well. Both drawers already
render it, so this covers them and any future drawer without a selector
list to keep in sync.

Two things the follow-up note didn't anticipate:

- The hook now reports { cards, modal } separately. A card overlap still
  lets the kitten wake for a bubble; a modal overlap yields
  unconditionally — a cheer isn't worth covering the panel she just
  opened on purpose.
- The speech bubble is its own layer, so fading the badge didn't hide it.
  Hold it back while a panel is open; useCompanion keeps it in state, so
  it reappears when she closes the panel.

Also guard the poll's setState on value equality, so the 500 ms tick
stops re-rendering the companion for an unchanged answer.

UX_REVIEW item 1 (redo does not re-apply an accepted suggestion) is
recorded as NOT REPRODUCIBLE. Read the prosemirror-history state directly
and hooked view.dispatch: redo works pressed immediately, after an 18 s
pause that lets a full re-check land, and with the editor never focused.
The doc's hypothesis is false — every re-check transaction is
decoration-only, which prosemirror-history ignores, and canRedo stayed
true throughout. Two real findings from that dig are written into the doc
instead: keyboard undo dies when focus isn't in the editor, and an undone
suggestion stays accepted server-side so its card doesn't reliably return.

Item 5's premise is also partly wrong and now re-scoped: the Chinese
sentence does produce a card with an English rendering, just labeled
Clarity rather than a first-class Translate type.

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 22:12:30 -07:00
prosolis ec9fba9252 Kitten yields to cards: fade, shrink, click-through; plus UX review doc
When a suggestion card drifts into the mascot's corner, the kitten turns
translucent (15%), steps back 10% (standalone `scale` so it composes with
the bob animation), and lets clicks pass through to the card. It wakes
while its bubble or the picker is open, or once the corner clears.

Also adds UX_REVIEW_2026-07-27.md — the hands-on review of the live deploy
turned into implementation-ready items (repro, location, fix, acceptance).

Claude-Session: https://claude.ai/code/session_016y6gyuHkQXPiEuW8RGQyua
2026-07-27 21:57:20 -07:00
88 changed files with 8295 additions and 336 deletions
+99 -3
View File
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
+3
View File
@@ -207,6 +207,9 @@ func main() {
lex := lexicon.NewHandler(database.DB, lexSet)
pr.Mount("/word", lex.Routes())
pr.Mount("/gloss", lex.GlossRoutes())
// The same lookup pointing the other way: a Chinese word to its pinyin
// and English senses, for an account whose direction is learning_pair.
pr.Mount("/hanzi", lex.HanziRoutes())
// Vocabulary garden: words the writer looks up are captured here and
// surfaced for gentle spaced-repetition review.
+3
View File
@@ -59,6 +59,9 @@ TTS_VOICE_EN=en_US-amy-medium
TTS_VOICE_ZH=zh_CN-huayan-medium
TTS_VOICE_PT=pt_PT-tugão-medium
TTS_VOICE_FR=fr_FR-siwis-medium
# Mexican, not peninsular — the es pack is written in neutral Latin American
# Spanish, and es_ES-davefx-medium would read it in the accent it avoids.
TTS_VOICE_ES=es_MX-ald-medium
TTS_AUDIO_FORMAT=mp3
TTS_TIMEOUT=15s
+24
View File
@@ -52,6 +52,7 @@ services:
# hyphen, and there is one Portuguese voice loaded either way.
TTS_ENDPOINT_PT: http://piper-pt:5000
TTS_ENDPOINT_FR: http://piper-fr:5000
TTS_ENDPOINT_ES: http://piper-es:5000
# The sidecars run piper-tts 1.6.0, which serves synthesis on
# /synthesize; millenia's older server keeps the default "/".
TTS_PATH: /synthesize
@@ -173,6 +174,29 @@ services:
networks:
- internal
# Spanish, for the es pair — and the Portuguese trap rather than the French
# one. Piper's catalogue has nine Spanish voices, six of them es_ES, and the
# obvious pick (es_ES-davefx-medium, which the build plan itself named) is
# peninsular. The es pack is written in neutral Latin American Spanish, so a
# Castilian voice would read it aloud in the accent the copy was written to
# avoid — the same wrong-country default that pt-PT hit through packaging,
# arriving here through the voice list. Only two Latin American voices exist,
# es_AR-daniela-high and es_MX; Mexican is the neutral broadcast standard and
# ald-medium matches the register of the other four. ASCII, so the
# percent-encoded download fallback added for tugão never has to fire.
piper-es:
build:
context: deploy/piper
image: petal-piper:local
container_name: petal-piper-es
restart: unless-stopped
environment:
PIPER_VOICE: ${TTS_VOICE_ES:-es_MX-ald-medium}
volumes:
- piper-voices:/voices
networks:
- internal
networks:
# Created and owned by the host's Traefik stack.
traefik:
+134 -11
View File
@@ -24,7 +24,7 @@ func patchMe(t *testing.T, users *UserStore, id, body string) *httptest.Response
func TestSetPairLang(t *testing.T) {
_, users, _ := newStores(t)
if err := users.SetPairLang("bob", "pt-PT"); err != nil {
if err := users.SetPair("bob", "pt-PT", DirectionLearningEn); err != nil {
t.Fatalf("set pt-PT: %v", err)
}
if u, _ := users.Get("bob"); u.PairLang != "pt-PT" {
@@ -34,16 +34,23 @@ func TestSetPairLang(t *testing.T) {
// Every pair with a langpack, not just the first one: this list and the
// frontend's PACKS are two copies of the same fact, and the day they
// disagree is the day she can pick a pair the app cannot render.
if err := users.SetPairLang("bob", "fr"); err != nil {
if err := users.SetPair("bob", "fr", DirectionLearningEn); err != nil {
t.Fatalf("set fr: %v", err)
}
if u, _ := users.Get("bob"); u.PairLang != "fr" {
t.Fatalf("pair_lang = %q, want fr", u.PairLang)
}
if err := users.SetPair("bob", "es", DirectionLearningEn); err != nil {
t.Fatalf("set es: %v", err)
}
if u, _ := users.Get("bob"); u.PairLang != "es" {
t.Fatalf("pair_lang = %q, want es", u.PairLang)
}
// And back — a writer who tries a pair and doesn't like it must be able to
// return, which is the whole reason the picker exists.
if err := users.SetPairLang("bob", "zh"); err != nil {
if err := users.SetPair("bob", "zh", DirectionLearningEn); err != nil {
t.Fatalf("set zh: %v", err)
}
if u, _ := users.Get("bob"); u.PairLang != "zh" {
@@ -56,11 +63,14 @@ func TestSetPairLang(t *testing.T) {
func TestSetPairLangRejectsUnshippedPairs(t *testing.T) {
_, users, _ := newStores(t)
// "es" is the real case here — the pair whose pack has not been written yet.
// "pt-BR" is the near-miss that matters most: a Brazilian code must not be
// quietly served European copy and a European voice.
for _, lang := range []string{"es", "pt-BR", "fr-CA", "klingon", "", " "} {
if err := users.SetPairLang("bob", lang); err == nil {
// The near-misses are the ones that matter, and there are two of them now.
// "pt-BR" must not be quietly served European copy and a European voice;
// "es-ES" is the same mistake pointing the other way, because the es pack is
// deliberately Latin American and reads itself aloud in a Mexican voice. A
// regional code Petal has not decided about is refused rather than rounded
// to the nearest pack it happens to have.
for _, lang := range []string{"es-ES", "pt-BR", "fr-CA", "de", "klingon", "", " "} {
if err := users.SetPair("bob", lang, DirectionLearningEn); err == nil {
t.Fatalf("stored unshipped pair %q", lang)
}
}
@@ -71,7 +81,7 @@ func TestSetPairLangRejectsUnshippedPairs(t *testing.T) {
func TestSetPairLangUnknownUser(t *testing.T) {
_, users, _ := newStores(t)
if err := users.SetPairLang("nobody", "pt-PT"); err == nil {
if err := users.SetPair("nobody", "pt-PT", DirectionLearningEn); err == nil {
t.Fatal("set a pair language on an account that does not exist")
}
}
@@ -98,8 +108,8 @@ func TestUpdateMeHandlerRejects(t *testing.T) {
_, users, _ := newStores(t)
for name, body := range map[string]string{
"unshipped pair": `{"pair_lang":"es"}`,
"missing field": `{}`,
"unshipped pair": `{"pair_lang":"es-ES"}`,
"unknown direction": `{"direction":"learning_klingon"}`,
"not json": `pt-PT`,
} {
if w := patchMe(t, users, "bob", body); w.Code != http.StatusBadRequest {
@@ -116,3 +126,116 @@ func TestUpdateMeHandlerRejects(t *testing.T) {
t.Fatalf("unknown user: status = %d, want 401", w.Code)
}
}
// An empty body used to be a 400, back when pair_lang was the only field and a
// request that named none of it could only be a client bug. With two optional
// fields it is an ordinary PATCH that changes nothing, and it has to be: the
// picker sends one field without knowing the other, and "omitted" has to mean
// "leave it alone" for that to be safe.
func TestUpdateMeHandlerEmptyBodyChangesNothing(t *testing.T) {
_, users, _ := newStores(t)
if err := users.SetPair("bob", "zh", DirectionLearningPair); err != nil {
t.Fatalf("set up: %v", err)
}
w := patchMe(t, users, "bob", `{}`)
if w.Code != http.StatusOK {
t.Fatalf("status = %d, want 200 (%s)", w.Code, w.Body.String())
}
u, _ := users.Get("bob")
if u.PairLang != "zh" || u.Direction != DirectionLearningPair {
t.Fatalf("empty PATCH moved the account to %q/%q", u.PairLang, u.Direction)
}
}
// The direction axis: an account can be turned around and turned back, and the
// default every existing row already carries is the one it had before the column
// existed.
func TestDirectionRoundTrip(t *testing.T) {
_, users, _ := newStores(t)
if u, _ := users.Get("bob"); u.Direction != DirectionLearningEn {
t.Fatalf("a fresh account starts at %q, want %q", u.Direction, DirectionLearningEn)
}
w := patchMe(t, users, "bob", `{"direction":"learning_pair"}`)
if w.Code != http.StatusOK {
t.Fatalf("turn around: status = %d (%s)", w.Code, w.Body.String())
}
var got db.User
if err := json.Unmarshal(w.Body.Bytes(), &got); err != nil {
t.Fatalf("decode: %v", err)
}
// The response carries the direction, not just the pair — the client reads
// its whole state back from here rather than assuming the write took.
if got.Direction != DirectionLearningPair || got.PairLang != "zh" {
t.Fatalf("response = %+v, want bob learning zh", got)
}
if w := patchMe(t, users, "bob", `{"direction":"learning_en"}`); w.Code != http.StatusOK {
t.Fatalf("turn back: status = %d (%s)", w.Code, w.Body.String())
}
if u, _ := users.Get("bob"); u.Direction != DirectionLearningEn {
t.Fatalf("direction = %q after turning back", u.Direction)
}
}
// The refusal this axis exists to make: a pair with no word list cannot be
// learned toward, however good its langpack is. fr, es and pt-PT all have copy,
// voices and spelling dictionaries — and nothing that could segment a sentence
// or read from that language into English, which is what a learner needs.
func TestLearnerDirectionRefusedForPairsWithoutData(t *testing.T) {
_, users, _ := newStores(t)
for _, lang := range []string{"pt-PT", "fr", "es"} {
if err := users.SetPair("bob", lang, DirectionLearningEn); err != nil {
t.Fatalf("set %s: %v", lang, err)
}
w := patchMe(t, users, "bob", `{"direction":"learning_pair"}`)
if w.Code != http.StatusBadRequest {
t.Fatalf("%s: status = %d, want 400", lang, w.Code)
}
if u, _ := users.Get("bob"); u.Direction != DirectionLearningEn {
t.Fatalf("%s: a refused write still moved direction to %q", lang, u.Direction)
}
}
}
// The two-field combination the handler validates as one decision. An account
// already learning Chinese that asks only to change pair is asking for a state
// neither field names on its own — French with segmentation — and it must not
// arrive by leaving one field out.
func TestPairChangeCannotStrandTheLearnerDirection(t *testing.T) {
_, users, _ := newStores(t)
if err := users.SetPair("bob", "zh", DirectionLearningPair); err != nil {
t.Fatalf("set up: %v", err)
}
if w := patchMe(t, users, "bob", `{"pair_lang":"fr"}`); w.Code != http.StatusBadRequest {
t.Fatalf("status = %d, want 400", w.Code)
}
u, _ := users.Get("bob")
if u.PairLang != "zh" || u.Direction != DirectionLearningPair {
t.Fatalf("refused write left the account at %q/%q", u.PairLang, u.Direction)
}
// Naming both at once is how that move is actually made, and it works.
if w := patchMe(t, users, "bob", `{"pair_lang":"fr","direction":"learning_en"}`); w.Code != http.StatusOK {
t.Fatalf("both fields: status = %d (%s)", w.Code, w.Body.String())
}
if u, _ := users.Get("bob"); u.PairLang != "fr" || u.Direction != DirectionLearningEn {
t.Fatalf("account = %q/%q, want fr/learning_en", u.PairLang, u.Direction)
}
}
// The CHECK constraint is the last line, below the handler and below SetPair:
// a direction that reaches the column by any other route is still refused.
func TestDirectionCheckConstraint(t *testing.T) {
_, users, database := newStores(t)
if _, err := database.Exec(`UPDATE users SET direction = 'sideways' WHERE id = 'bob'`); err == nil {
t.Fatal("the users.direction CHECK accepted 'sideways'")
}
if u, _ := users.Get("bob"); u.Direction != DirectionLearningEn {
t.Fatalf("direction = %q after a refused UPDATE", u.Direction)
}
}
+112 -13
View File
@@ -49,9 +49,9 @@ func (u *UserStore) Upsert(sub, email, displayName string) error {
func (u *UserStore) Get(id string) (db.User, error) {
var user db.User
err := u.db.QueryRow(
`SELECT id, email, COALESCE(display_name, ''), created_at, pair_lang
`SELECT id, email, COALESCE(display_name, ''), created_at, pair_lang, direction
FROM users WHERE id = ?`, id,
).Scan(&user.ID, &user.Email, &user.DisplayName, &user.CreatedAt, &user.PairLang)
).Scan(&user.ID, &user.Email, &user.DisplayName, &user.CreatedAt, &user.PairLang, &user.Direction)
return user, err
}
@@ -76,8 +76,13 @@ func (u *UserStore) MeHandler() http.HandlerFunc {
// this one names the pairs Petal can render itself in, which requires a langpack
// on the frontend. Accepting a code with no pack would leave her looking at
// Chinese with no way back except another guess, so the server refuses it. es
// joins this list on the day its pack lands, not before.
var shippedPairs = []string{"zh", "pt-PT", "fr"}
// joined on the day its pack landed, not before.
//
// These four are now every pair PairLang names on the frontend, which makes the
// two lists look redundant. They are not: the next pair will exist in the type
// and in the prompts long before it has copy, and this list is the one that
// says a writer may actually be sent there.
var shippedPairs = []string{"zh", "pt-PT", "fr", "es"}
func pairIsShipped(lang string) bool {
for _, p := range shippedPairs {
@@ -88,12 +93,63 @@ func pairIsShipped(lang string) bool {
return false
}
// SetPairLang moves an account to another (English + X) pair.
func (u *UserStore) SetPairLang(id, lang string) error {
// The two directions a pair can be travelled in. `DirectionLearningEn` is the
// original assumption made explicit: the writer is native in X and practising
// English. `DirectionLearningPair` is the other way round.
const (
DirectionLearningEn = "learning_en"
DirectionLearningPair = "learning_pair"
)
// The pairs whose *learner* direction Petal can actually serve, which is a
// narrower thing than a shipped pair and narrower again than a langpack.
//
// Turning a pair around needs data no langpack carries: a word list to segment
// with, and a dictionary that reads from the pair language into English. Chinese
// has both as of Phase 26 (CC-CEDICT + jieba); French, Spanish and Portuguese
// have neither yet, and — unlike a missing pack, which leaves a writer looking
// at copy she cannot read — a missing word list would leave her looking at an
// editor that silently does nothing when she hovers. Both are bad; only one is
// legible as a bug. So the server refuses, for the same reason and by the same
// mechanism as `shippedPairs`.
//
// This list is expected to grow one pair at a time and never to be inferred:
// segmentation is a property of a writing system, and there is no rule that
// derives "has a word list" from a language code.
var learnerPairs = []string{"zh"}
// SupportsLearnerDirection reports whether a pair can be turned around.
func SupportsLearnerDirection(lang string) bool {
for _, p := range learnerPairs {
if p == lang {
return true
}
}
return false
}
func directionIsKnown(d string) bool {
return d == DirectionLearningEn || d == DirectionLearningPair
}
// SetPair moves an account to another (English + X) pair, in a given direction.
//
// The two are written together because they constrain each other: a direction is
// only meaningful for a pair that can be travelled in it, and validating them a
// field at a time would let a two-step change pass through a state that neither
// step is allowed to leave behind.
func (u *UserStore) SetPair(id, lang, direction string) error {
if !pairIsShipped(lang) {
return errors.New("auth: unshipped pair language " + lang)
}
res, err := u.db.Exec(`UPDATE users SET pair_lang = ? WHERE id = ?`, lang, id)
if !directionIsKnown(direction) {
return errors.New("auth: unknown direction " + direction)
}
if direction == DirectionLearningPair && !SupportsLearnerDirection(lang) {
return errors.New("auth: no learner direction for " + lang)
}
res, err := u.db.Exec(
`UPDATE users SET pair_lang = ?, direction = ? WHERE id = ?`, lang, direction, id)
if err != nil {
return err
}
@@ -103,8 +159,8 @@ func (u *UserStore) SetPairLang(id, lang string) error {
return nil
}
// UpdateMeHandler changes the caller's own settings — today, the one setting
// there is: which language Petal speaks alongside her English.
// UpdateMeHandler changes the caller's own settings: which language Petal
// speaks alongside her English, and which of the two she is learning.
//
// It answers with the whole updated user rather than an empty 204 so the client
// has one shape to trust: /api/me and this return the same thing, and the app
@@ -114,24 +170,67 @@ func (u *UserStore) SetPairLang(id, lang string) error {
// dictionary, her read-aloud voice, which word-lookup provider answers, and the
// language the prompts ask the model to explain in. All of those read
// `users.pair_lang` at use time, so all of them follow from this one write.
//
// Both fields are optional and each defaults to what the account already has, so
// the picker can send one without knowing the other. That matters for the
// combination this endpoint exists to prevent: a client that sent only
// `pair_lang: "fr"` while the account sat on `learning_pair` would otherwise ask
// for French-with-segmentation, which does not exist. Here it is one decision
// with one validation, and the answer carries whatever actually landed.
func (u *UserStore) UpdateMeHandler() http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
var body struct {
PairLang string `json:"pair_lang"`
PairLang *string `json:"pair_lang"`
Direction *string `json:"direction"`
}
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
httputil.BadRequest(w, "invalid request body")
return
}
lang := strings.TrimSpace(body.PairLang)
id := UserID(r.Context())
current, err := u.Get(id)
if err != nil {
// Only a missing row means "not signed in". A dictionary-file or
// SQLite fault answered as 401 would trip the client's session
// interceptor and throw a writer out of an app she is still signed
// in to — the same distinction SetPair's error branch makes below.
if errors.Is(err, sql.ErrNoRows) {
httputil.ErrorJSON(w, http.StatusUnauthorized, "not signed in")
return
}
httputil.ServerError(w, err)
return
}
lang, direction := current.PairLang, current.Direction
if body.PairLang != nil {
lang = strings.TrimSpace(*body.PairLang)
}
if body.Direction != nil {
direction = strings.TrimSpace(*body.Direction)
}
if !pairIsShipped(lang) {
// Name the ones that work. A writer who lands here has picked from a
// stale client, and "not a language" tells her nothing.
httputil.BadRequest(w, "unsupported language pair — Petal speaks "+strings.Join(shippedPairs, ", "))
return
}
id := UserID(r.Context())
if err := u.SetPairLang(id, lang); err != nil {
if !directionIsKnown(direction) {
httputil.BadRequest(w, "unknown direction — expected "+DirectionLearningEn+" or "+DirectionLearningPair)
return
}
if direction == DirectionLearningPair && !SupportsLearnerDirection(lang) {
// Refused rather than quietly downgraded to learning_en. A silent
// downgrade would leave the writer looking at an editor that behaves
// like the one she just tried to leave, with nothing to read as an
// explanation — and the caller cannot tell the two outcomes apart
// without diffing the response it was given.
httputil.BadRequest(w, "Petal can only be learned toward "+strings.Join(learnerPairs, ", ")+" so far")
return
}
if err := u.SetPair(id, lang, direction); err != nil {
if errors.Is(err, sql.ErrNoRows) {
httputil.ErrorJSON(w, http.StatusUnauthorized, "not signed in")
return
+6
View File
@@ -16,6 +16,8 @@ func TestTTSVoicesDiscovery(t *testing.T) {
"TTS_VOICE_PT=pt_PT-tugão-medium",
"TTS_ENDPOINT_FR=http://piper-fr:5000",
"TTS_VOICE_FR=fr_FR-siwis-medium",
"TTS_ENDPOINT_ES=http://piper-es:5000",
"TTS_VOICE_ES=es_MX-ald-medium",
// Noise that must not become a language.
"TTS_PATH=/synthesize",
"PATH=/usr/bin",
@@ -30,6 +32,10 @@ func TestTTSVoicesDiscovery(t *testing.T) {
// Phase 24's whole TTS change: a fourth language costs two lines here
// and a compose service, and no Go at all.
"fr": {"http://piper-fr:5000", "fr_FR-siwis-medium"},
// And a fifth cost exactly the same, which is the claim actually being
// tested. The voice is Mexican on purpose: the es pack is Latin
// American, and es_ES-davefx-medium would read it in the wrong accent.
"es": {"http://piper-es:5000", "es_MX-ald-medium"},
}
if len(voices) != len(want) {
t.Fatalf("discovered %v, want %v", voices, want)
+92
View File
@@ -497,6 +497,98 @@ CREATE INDEX idx_suggestions_resolved ON suggestions(status, resolved_at);
stmt: `
ALTER TABLE suggestions ADD COLUMN source TEXT NOT NULL DEFAULT 'llm';
UPDATE suggestions SET source = 'local' WHERE type = 'mechanics';
`,
},
{
// Sentence-level identity, so a re-check stops regenerating the world.
// Every pass used to delete its whole family and re-insert it, which
// meant accepting one edit gave every other card a new id and a newly
// worded explanation — the rail visibly emptied and refilled, and the
// model was asked again about sentences nobody had touched.
//
// `chunk_hash` records which sentence a suggestion belongs to, and
// checked_chunks records which sentences a family has already read. A
// re-check then asks only about the difference and keeps the rest of
// the rows exactly as they are, id and wording included.
//
// Existing rows get '' — "sentence unknown", which reads as in-play, so
// they are simply reconciled on the next pass like any fresh finding.
name: "0014_suggestion_chunk_hash",
stmt: `
ALTER TABLE suggestions ADD COLUMN chunk_hash TEXT NOT NULL DEFAULT '';
CREATE TABLE checked_chunks (
doc_id TEXT NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
family TEXT NOT NULL,
hash TEXT NOT NULL,
PRIMARY KEY (doc_id, family, hash)
);
`,
},
{
// A sentence she wrote in her own language gets its own type. Petal already
// detected such spans and already rendered them into English — it just
// filed the result under 'clarity', so the pair model's flagship moment
// read as tidying up her Chinese. As with 0005 and 0008, the `type` CHECK
// can't be ALTERed in place, so rebuild the table with the extended
// constraint, copy every row across, and recreate both indexes.
//
// Existing rows are left on whatever type they have. A card she is already
// reading keeps the label she has already read (the same rule reconcile.go
// follows for a re-proposed edit); new findings get the new label.
name: "0015_translate_suggestion_type",
stmt: `
CREATE TABLE suggestions_new (
id TEXT PRIMARY KEY DEFAULT (lower(hex(randomblob(16)))),
doc_id TEXT NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
from_pos INTEGER NOT NULL,
to_pos INTEGER NOT NULL,
original TEXT NOT NULL,
replacement TEXT NOT NULL,
explanation TEXT NOT NULL,
type TEXT NOT NULL CHECK(type IN ('grammar','phrasing','idiom','clarity','translate','voice','collocation','mechanics')),
status TEXT NOT NULL DEFAULT 'pending' CHECK(status IN ('pending','accepted','rejected')),
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
resolved_at DATETIME,
source TEXT NOT NULL DEFAULT 'llm',
chunk_hash TEXT NOT NULL DEFAULT ''
);
INSERT INTO suggestions_new (id, doc_id, from_pos, to_pos, original, replacement, explanation, type, status, created_at, resolved_at, source, chunk_hash)
SELECT id, doc_id, from_pos, to_pos, original, replacement, explanation, type, status, created_at, resolved_at, source, chunk_hash FROM suggestions;
DROP TABLE suggestions;
ALTER TABLE suggestions_new RENAME TO suggestions;
CREATE INDEX idx_suggestions_doc_id ON suggestions(doc_id);
CREATE INDEX idx_suggestions_resolved ON suggestions(status, resolved_at);
`,
},
{
// Which half of the pair is being learned.
//
// `pair_lang` (0010) has always answered "which two languages", and every
// surface built on it assumed the answer to a second question nobody had
// asked: that English is the language being *learned*. That assumption is
// load-bearing in a dozen places — CJK is deliberately never tokenized,
// never spell-checked, never glossed; the prompts explain English in her
// language; the vocabulary garden captures English words. All correct for
// a Mandarin native practising English, and all backwards for an English
// native practising Mandarin.
//
// A second pair code ('zh-learner') was the cheaper option and is the
// wrong shape: it would make the two directions of one pair look like two
// unrelated languages to every query, and it would have to be repeated for
// fr, es and pt-PT before any of them could turn around. A column keeps
// the two questions separate, which is what they are.
//
// 'learning_en' is the default and is what every existing row means — the
// backfill is the DEFAULT itself, and it is right rather than merely
// convenient: all three accounts today are Mandarin natives writing
// English.
name: "0016_user_direction",
stmt: `
ALTER TABLE users ADD COLUMN direction TEXT NOT NULL DEFAULT 'learning_en'
CHECK(direction IN ('learning_en','learning_pair'));
`,
},
}
+137
View File
@@ -3,6 +3,7 @@ package db
import (
"path/filepath"
"testing"
"time"
)
func TestOpenMigratesAndSeeds(t *testing.T) {
@@ -227,3 +228,139 @@ func TestSuggestionSourceBackfill(t *testing.T) {
t.Errorf("default source = %q, want %q", fresh, SuggestionSourceLLM)
}
}
// TestTranslateTypeMigrationPreservesRows runs migration 0015 against a database
// that predates it. Unlike the two backfills above, 0015 *rebuilds the table* —
// SQLite can't ALTER a CHECK constraint — so it copies every row across by hand,
// and a column left out of that copy list silently loses her data. Every test
// elsewhere starts from a fresh database and would never notice; the live box has
// years of rows in it.
func TestTranslateTypeMigrationPreservesRows(t *testing.T) {
path := filepath.Join(t.TempDir(), "old.db")
d, err := Open(path)
if err != nil {
t.Fatalf("open: %v", err)
}
// Rewind to the pre-0015 table: the same shape, minus 'translate' in the CHECK.
if _, err := d.Exec(`
CREATE TABLE suggestions_old (
id TEXT PRIMARY KEY DEFAULT (lower(hex(randomblob(16)))),
doc_id TEXT NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
from_pos INTEGER NOT NULL,
to_pos INTEGER NOT NULL,
original TEXT NOT NULL,
replacement TEXT NOT NULL,
explanation TEXT NOT NULL,
type TEXT NOT NULL CHECK(type IN ('grammar','phrasing','idiom','clarity','voice','collocation','mechanics')),
status TEXT NOT NULL DEFAULT 'pending' CHECK(status IN ('pending','accepted','rejected')),
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
resolved_at DATETIME,
source TEXT NOT NULL DEFAULT 'llm',
chunk_hash TEXT NOT NULL DEFAULT ''
);
DROP TABLE suggestions;
ALTER TABLE suggestions_old RENAME TO suggestions;
CREATE INDEX idx_suggestions_doc_id ON suggestions(doc_id);
CREATE INDEX idx_suggestions_resolved ON suggestions(status, resolved_at);
DELETE FROM schema_migrations WHERE name = '0015_translate_suggestion_type';
`); err != nil {
t.Fatalf("rewind schema: %v", err)
}
if _, err := d.Exec(`INSERT INTO documents (id, user_id) VALUES ('d1', ?)`, LocalUserID); err != nil {
t.Fatalf("insert document: %v", err)
}
// One row with every column carrying a distinguishable value, so a dropped
// column shows up as a changed value rather than as a passing test.
if _, err := d.Exec(
`INSERT INTO suggestions (id, doc_id, from_pos, to_pos, original, replacement, explanation, type, status, created_at, resolved_at, source, chunk_hash)
VALUES ('s-1', 'd1', 7, 11, 'by foots', 'on foot', 'idiom advice she has read', 'idiom', 'accepted', '2026-01-02 03:04:05', '2026-01-02 03:05:00', 'local', 'abc123')`,
); err != nil {
t.Fatalf("seed row: %v", err)
}
d.Close()
d2, err := Open(path)
if err != nil {
t.Fatalf("reopen (migrate): %v", err)
}
defer d2.Close()
var (
docID, original, replacement, explanation string
typ, status, source, chunkHash string
from, to int
// Scanned as instants, not strings: the driver renders a DATETIME column in
// its own format, so the claim is "the same moment", not the same text.
createdAt, resolvedAt time.Time
)
if err := d2.QueryRow(
`SELECT doc_id, from_pos, to_pos, original, replacement, explanation, type, status, created_at, resolved_at, source, chunk_hash
FROM suggestions WHERE id = 's-1'`,
).Scan(&docID, &from, &to, &original, &replacement, &explanation,
&typ, &status, &createdAt, &resolvedAt, &source, &chunkHash); err != nil {
t.Fatalf("read migrated row: %v", err)
}
for _, c := range []struct{ name, got, want string }{
{"doc_id", docID, "d1"},
{"original", original, "by foots"},
{"replacement", replacement, "on foot"},
{"explanation", explanation, "idiom advice she has read"},
{"type", typ, SuggestionTypeIdiom},
{"status", status, SuggestionStatusAccepted},
{"source", source, SuggestionSourceLocal},
{"chunk_hash", chunkHash, "abc123"},
} {
if c.got != c.want {
t.Errorf("%s = %q, want %q", c.name, c.got, c.want)
}
}
if from != 7 || to != 11 {
t.Errorf("offsets = (%d, %d), want (7, 11)", from, to)
}
// created_at and resolved_at must survive: the rail's arrival chime keys on
// created_at, and the growth journal counts by resolved_at. A rebuild that
// reset either would re-chime her whole document and rewrite her history.
for _, c := range []struct {
name string
got time.Time
want string
}{
{"created_at", createdAt, "2026-01-02 03:04:05"},
{"resolved_at", resolvedAt, "2026-01-02 03:05:00"},
} {
want, err := time.Parse("2006-01-02 15:04:05", c.want)
if err != nil {
t.Fatalf("parse want: %v", err)
}
if !c.got.Equal(want) {
t.Errorf("%s = %v, want the original instant %v", c.name, c.got, want)
}
}
// The point of the rebuild: the new type is now insertable, and a bogus one
// still isn't.
if _, err := d2.Exec(
`INSERT INTO suggestions (id, doc_id, from_pos, to_pos, original, replacement, explanation, type)
VALUES ('s-2', 'd1', 0, 3, '苹果', 'apple', 'x', ?)`, SuggestionTypeTranslate,
); err != nil {
t.Fatalf("insert translate row: %v", err)
}
if _, err := d2.Exec(
`INSERT INTO suggestions (id, doc_id, from_pos, to_pos, original, replacement, explanation, type)
VALUES ('s-3', 'd1', 0, 3, 'x', 'y', 'x', 'nonsense')`,
); err == nil {
t.Error("CHECK constraint accepted an unknown type after the rebuild")
}
// Both indexes must come back, or every document load starts table-scanning.
for _, idx := range []string{"idx_suggestions_doc_id", "idx_suggestions_resolved"} {
var name string
if err := d2.QueryRow(
`SELECT name FROM sqlite_master WHERE type = 'index' AND name = ?`, idx,
).Scan(&name); err != nil {
t.Errorf("index %s missing after rebuild: %v", idx, err)
}
}
}
+21 -1
View File
@@ -15,6 +15,19 @@ type User struct {
// today, "pt-PT"/"fr"/"es" once the langpacks land. It selects the UI copy
// and dictionary set, not the language they may type in.
PairLang string `json:"pair_lang"`
// Direction says which half of the pair is being *learned*. Every pair until
// now assumed one answer: the writer is native in X and practising English,
// so hanzi is never tokenized and English is what gets underlined. Turn it
// around — a native English speaker learning Chinese — and the same pair
// wants the opposite of nearly every default.
//
// It is a separate column from PairLang rather than a second pair code
// ("zh-learner") because it is a genuinely separate question: the pair says
// *which two languages*, this says *which way round*. Keeping them apart is
// what lets fr, es and pt-PT inherit the learner direction later without a
// second langpack each.
Direction string `json:"direction"`
}
// Document is a single piece of writing. `Content` is the Tiptap JSON document
@@ -104,7 +117,7 @@ type Suggestion struct {
Original string `json:"original"`
Replacement string `json:"replacement"`
Explanation string `json:"explanation"`
Type string `json:"type"` // grammar | phrasing | idiom | clarity | voice | collocation
Type string `json:"type"` // grammar | phrasing | idiom | clarity | translate | voice | collocation
Status string `json:"status"` // pending | accepted | rejected
// Source names the engine that proposed the edit, not its family: an offline
// rule and the model can both propose a collocation, and the writer is never
@@ -119,6 +132,13 @@ const (
SuggestionTypePhrasing = "phrasing"
SuggestionTypeIdiom = "idiom"
SuggestionTypeClarity = "clarity"
// A span she wrote in her own language, rendered into English. Not a
// correction — nothing was wrong with it — which is why it is its own type
// rather than a clarity fix: the card is the pair model's flagship moment
// (SUGGESTIONS §1), and labelling it "Clarity" reads as a tidy-up of her
// first language. The model isn't asked for this label; it is derived from the
// span itself (see suggestions/language.go), so it can't drift.
SuggestionTypeTranslate = "translate"
SuggestionTypeVoice = "voice"
SuggestionTypeCollocation = "collocation"
SuggestionTypeMechanics = "mechanics" // deterministic rule-based pass (no LLM)
+10
View File
@@ -36,3 +36,13 @@ var glossGz []byte
//
//go:embed data/phonetic.json.gz
var phoneticGz []byte
// hanziGz is the gzipped Chinese→English map: simplified headword → [[pinyin,
// senses], …]. Built from CC-CEDICT (scripts/build_cedict.py), unfiltered — the
// word a learner stops on is the one they do not know, so this is the one
// dataset here with no frequency gate. Loaded on its own sync.Once (see
// hanzi.go), not with the four above, because only a learner-direction account
// ever asks for it.
//
//go:embed data/hanzi.json.gz
var hanziGz []byte
Binary file not shown.
+30
View File
@@ -42,6 +42,36 @@ func (h *Handler) GlossRoutes() chi.Router {
return r
}
// HanziRoutes returns the router mounted at /api/hanzi — a Chinese word to its
// pinyin and English senses, for a writer going the other way through the zh
// pair (`users.direction = 'learning_pair'`).
//
// It does not go through [Handler.providerFor], and that is not an oversight.
// providerFor picks a dictionary by the writer's *pair*, to answer "what does
// this English word mean in her language" — a question whose answer differs per
// pair. This endpoint asks the opposite question of exactly one language, and
// [auth.SupportsLearnerDirection] already guarantees that language is Chinese.
// Routing it through the pair would add a database read per hover to choose
// between one option and itself.
func (h *Handler) HanziRoutes() chi.Router {
r := chi.NewRouter()
r.Get("/{word}", h.hanzi)
return r
}
// hanzi answers a Chinese word lookup. Like the other two, a miss is a 200 with
// empty lists — a hover that lands on a word the dictionary has never heard of
// is an ordinary thing to happen while reading, and the tooltip simply doesn't
// open.
func (h *Handler) hanzi(w http.ResponseWriter, r *http.Request) {
res, err := h.Set.Hanzi(pathWord(r))
if err != nil {
writeLookupErr(w, err)
return
}
writeLookup(w, res)
}
// providerFor returns the provider for the caller's language pair.
//
// The pair language is read here rather than threaded down because a word
+144
View File
@@ -0,0 +1,144 @@
package lexicon
import (
"fmt"
"strings"
"sync"
"unicode"
)
// The Chinese half of the lexicon: a word written in hanzi to its pinyin and
// English senses. This is the mirror image of `gloss` — that one reads English
// and answers in Chinese, for a Mandarin native practising English; this one
// reads Chinese and answers in English, for the other direction of the same
// pair (`users.direction = 'learning_pair'`).
//
// It is deliberately not folded into [Lexicon.load]. That method reads four
// datasets on the first lookup of any kind, and this one is 3.1 MB gzipped that
// only a learner-direction account will ever ask for — every other writer would
// pay the decompression and the resident memory for a map they never touch. Its
// own sync.Once means the cost lands on the first Chinese hover and nowhere
// else.
// HanziReading is one pronunciation of a word and the senses it carries in that
// pronunciation. A word usually has one; the ones that have two are why this is
// a list rather than a pair of strings. 得 is dé, "to obtain", *and* de, the
// particle that makes 说得很好 mean "speaks well" — a learner shown only the
// first has been told something false about the sentence in front of them.
type HanziReading struct {
Pinyin string `json:"pinyin"`
Senses string `json:"senses"`
}
// HanziChar is one character of a word that the dictionary could not answer as
// a whole. See [Lexicon.Hanzi].
type HanziChar struct {
Char string `json:"char"`
Pinyin string `json:"pinyin"`
Senses string `json:"senses"`
}
// HanziResult is what a Chinese word lookup answers. Readings is empty for a
// word the dictionary does not have, in which case Chars may carry the
// character-by-character reading instead.
type HanziResult struct {
Word string `json:"word"`
Readings []HanziReading `json:"readings"`
Chars []HanziChar `json:"chars"`
}
type hanziStore struct {
once sync.Once
err error
// word → [[pinyin, senses], …], exactly as scripts/build_cedict.py writes it.
entries map[string][][]string
}
var hanzi hanziStore
func (h *hanziStore) load() {
h.once.Do(func() {
if err := gunzipJSON(hanziGz, &h.entries); err != nil {
h.err = fmt.Errorf("load hanzi: %w", err)
}
})
}
// maxHanziChars caps the per-character fallback. A run longer than this is
// almost certainly a phrase the segmenter split badly rather than a word, and
// spelling out eight characters one at a time is a wall, not a hint.
const maxHanziChars = 6
// Hanzi returns the pinyin and English senses of a Chinese word.
//
// There is no de-inflection walk here, and its absence is a fact about the
// language rather than an omission: Chinese words do not inflect, so the
// candidate forms [lookupGloss] tries for "running" → "run" have no analogue.
// A lookup either hits the headword or it does not.
//
// What it does instead is fall back to the characters. The segmentation word
// list is a superset of this dictionary — every glossable word can be
// segmented, but jieba knows ordinary compounds CC-CEDICT has no entry for — so
// a hover really can land on a word with nothing to say about it. Chinese
// compounds are usually transparent from their parts (电脑 is "electric brain"),
// which makes the character reading a genuinely useful second answer rather
// than a consolation prize. It is returned as its own field so the surface can
// say which of the two it is showing; a caller that only wants whole words can
// ignore it.
func (l *Lexicon) Hanzi(word string) (HanziResult, error) {
hanzi.load()
if hanzi.err != nil {
return HanziResult{}, hanzi.err
}
norm := strings.TrimSpace(word)
res := HanziResult{Word: word, Readings: []HanziReading{}, Chars: []HanziChar{}}
if norm == "" {
return res, nil
}
if rows, ok := hanzi.entries[norm]; ok {
res.Readings = toReadings(rows)
return res, nil
}
chars := []rune(norm)
if len(chars) < 2 || len(chars) > maxHanziChars {
// A single character that missed has no parts to fall back to, and a long
// run is not a word. Either way the honest answer is nothing.
return res, nil
}
for _, r := range chars {
if !unicode.Is(unicode.Han, r) {
// Mixed input (a stray letter or digit inside the run) is not something
// the character reading can explain, and guessing at the hanzi parts of
// it would be worse than silence.
return HanziResult{Word: word, Readings: []HanziReading{}, Chars: []HanziChar{}}, nil
}
rows, ok := hanzi.entries[string(r)]
if !ok {
continue
}
first := toReadings(rows)
if len(first) == 0 {
continue
}
res.Chars = append(res.Chars, HanziChar{
Char: string(r),
Pinyin: first[0].Pinyin,
Senses: first[0].Senses,
})
}
return res, nil
}
func toReadings(rows [][]string) []HanziReading {
out := make([]HanziReading, 0, len(rows))
for _, row := range rows {
if len(row) < 2 {
continue
}
out = append(out, HanziReading{Pinyin: row[0], Senses: row[1]})
}
return out
}
+150
View File
@@ -0,0 +1,150 @@
package lexicon
import (
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/go-chi/chi/v5"
)
// The Chinese direction of the lexicon, against the real embedded asset — not a
// fixture. The dataset is built by scripts/build_cedict.py, which asserts its
// own invariants at build time; what these assert is that the *lookup* over it
// behaves, including on the entries the build script goes out of its way to keep.
func TestHanziLookup(t *testing.T) {
l := New()
res, err := l.Hanzi("公园")
if err != nil {
t.Fatalf("lookup 公园: %v", err)
}
if len(res.Readings) == 0 {
t.Fatal("公园 has no readings")
}
// Tone marks, not the numbered pinyin CC-CEDICT stores. The number is the
// storage format; the marks are what a learner reads.
if got := res.Readings[0].Pinyin; got != "gōngyuán" {
t.Errorf("公园 pinyin = %q, want gōngyuán", got)
}
if !strings.Contains(res.Readings[0].Senses, "park") {
t.Errorf("公园 senses = %q, want something about a park", res.Readings[0].Senses)
}
// A word answered whole says nothing about its characters — the fallback is
// the other branch, and sending both would double the payload of the common
// case to no purpose.
if len(res.Chars) != 0 {
t.Errorf("a whole-word hit also returned %d characters", len(res.Chars))
}
}
// 得 is the reason readings are a list. Answered with only dé "to obtain", a
// learner hovering it in 说得很好 has been told something false about the
// sentence they are looking at.
func TestHanziParticleCarriesItsGrammaticalReading(t *testing.T) {
l := New()
for _, particle := range []string{"的", "地", "得"} {
res, err := l.Hanzi(particle)
if err != nil {
t.Fatalf("lookup %s: %v", particle, err)
}
var found bool
for _, r := range res.Readings {
if r.Pinyin == "de" {
found = true
}
}
if !found {
t.Errorf("%s never reads as neutral \"de\": %+v", particle, res.Readings)
}
}
}
// The fallback the segmentation gap makes necessary: jieba knows ordinary
// compounds CC-CEDICT has no headword for, so a hover can land on a real word
// with no entry. Chinese compounds are usually transparent from their parts, so
// the characters are a real second answer.
func TestHanziFallsBackToCharacters(t *testing.T) {
l := New()
// Constructed rather than borrowed from the corpus: a word that CC-CEDICT
// *does* carry would test the other branch, and which compounds it happens to
// omit is not something a test should pin.
const made = "猫书"
if _, ok := hanzi.entries[made]; ok {
t.Skipf("%s has become a real headword; pick another compound", made)
}
res, err := l.Hanzi(made)
if err != nil {
t.Fatalf("lookup %s: %v", made, err)
}
if len(res.Readings) != 0 {
t.Fatalf("%s answered as a whole word: %+v", made, res.Readings)
}
if len(res.Chars) != 2 {
t.Fatalf("character fallback gave %d entries, want 2: %+v", len(res.Chars), res.Chars)
}
if res.Chars[0].Char != "猫" || !strings.Contains(res.Chars[0].Senses, "cat") {
t.Errorf("first character = %+v, want 猫 ~ cat", res.Chars[0])
}
if res.Chars[0].Pinyin != "māo" {
t.Errorf("猫 pinyin = %q, want māo", res.Chars[0].Pinyin)
}
}
func TestHanziMisses(t *testing.T) {
l := New()
for name, word := range map[string]string{
// A single character with no entry has no parts to fall back to.
"lone unknown character": "龥",
"empty": "",
"whitespace": " ",
// Not Chinese at all: the English tokenizer owns these, and answering
// would mean guessing.
"english": "hello",
"mixed": "猫cat",
// Longer than a word: a bad segmentation, not something to spell out
// character by character.
"a whole clause": "我今天早上去公园跑步了",
} {
res, err := l.Hanzi(word)
if err != nil {
t.Fatalf("%s: %v", name, err)
}
if len(res.Readings) != 0 || len(res.Chars) != 0 {
t.Errorf("%s (%q) answered with %+v / %+v", name, word, res.Readings, res.Chars)
}
}
}
func TestHanziEndpoint(t *testing.T) {
h := NewHandler(nil, NewSet(nil))
r := chi.NewRouter()
r.Mount("/hanzi", h.HanziRoutes())
w := httptest.NewRecorder()
r.ServeHTTP(w, httptest.NewRequest(http.MethodGet, "/hanzi/"+"跑步", nil))
if w.Code != http.StatusOK {
t.Fatalf("status = %d", w.Code)
}
var got HanziResult
if err := json.Unmarshal(w.Body.Bytes(), &got); err != nil {
t.Fatalf("decode: %v", err)
}
if got.Word != "跑步" || len(got.Readings) == 0 || got.Readings[0].Pinyin != "pǎobù" {
t.Fatalf("response = %+v", got)
}
// A miss is a 200 with empty lists, like the other two lookups — the tooltip
// quietly doesn't open rather than showing an error over her writing.
w = httptest.NewRecorder()
r.ServeHTTP(w, httptest.NewRequest(http.MethodGet, "/hanzi/zzz", nil))
if w.Code != http.StatusOK {
t.Fatalf("miss: status = %d, want 200", w.Code)
}
}
+10
View File
@@ -109,3 +109,13 @@ func (g glossless) Lookup(word string) (Result, error) {
func (g glossless) Gloss(word string) (GlossResult, error) {
return GlossResult{Word: word}, nil
}
// Hanzi answers a Chinese-word lookup from the embedded CC-CEDICT map.
//
// It is on the Set rather than on [Provider] because it is not the same
// question the other two ask. Lookup and Gloss vary by pair — which is why they
// are behind an interface with two implementations — while this one is asked of
// Chinese or not at all: the learner direction exists for exactly one pair (see
// auth.learnerPairs), and DreamDict's own CC-CEDICT would be a second copy of
// the same dictionary, chosen by a rule with one branch.
func (s *Set) Hanzi(word string) (HanziResult, error) { return s.embedded.Hanzi(word) }
+37
View File
@@ -83,3 +83,40 @@ func TestDefaultPairStillReadsAsBefore(t *testing.T) {
t.Fatalf("zh ask-petal lost its Mandarin \"why\":\n%s", got)
}
}
// UX item 6: the Ask Petal answer is bilingual, pair language first, halves
// separated by one blank line. That separator is not a stylistic preference —
// AskPetal.tsx splits on it to render the two halves the way the companion
// renders its two lines — so the instruction has to survive prompt edits.
//
// The direction the writer is learning in is deliberately not encoded: the pair
// is (English + X), and an English speaker learning French needs the same two
// halves a Mandarin speaker learning English does. The prompt asks for both and
// lets the reader choose, so there is nothing here that names one half the
// answer and the other a courtesy.
func TestAskPetalAnswersInBothLanguages(t *testing.T) {
for _, code := range []string{"zh", "pt-PT", "fr", "es"} {
lang := LangFor(code)
ask := AskPetalSystemPrompt("a", "b", "grammar", "d", "e", lang)
if !strings.Contains(ask, "BOTH languages") {
t.Fatalf("%s: ask-petal no longer asks for both languages:\n%s", code, ask)
}
if !strings.Contains(ask, "single blank line") {
t.Fatalf("%s: ask-petal lost the blank-line separator the client splits on:\n%s", code, ask)
}
// Order matters to the rendering: the pair language is the prominent
// half, English the muted one beneath it.
if !strings.Contains(ask, "first the whole answer in "+lang.Name) {
t.Fatalf("%s: ask-petal doesn't put %s first:\n%s", code, lang.Name, ask)
}
// The instruction it replaced. Left in place it directly contradicts the
// new one, and a model given both will pick one at random.
if strings.Contains(ask, "Never mix languages") {
t.Fatalf("%s: ask-petal still forbids the bilingual reply it now asks for:\n%s", code, ask)
}
if strings.Contains(ask, "%!") {
t.Fatalf("%s: ask-petal prompt has a formatting error:\n%s", code, ask)
}
}
}
+29 -4
View File
@@ -144,6 +144,25 @@ func CollocationMessages(contentText, tone string, lang Lang) []Message {
// askPetalSystemTemplate is the Ask Petal tutor prompt. The suggestion context
// is interpolated in; the user's own messages are appended after this system
// turn by the caller.
//
// The reply is bilingual, the pair language first. Until UX item 6 it mirrored
// the language of the question instead — self-consistent, but it meant asking in
// one language cost you the other, and the writer doesn't always know which one
// the answer will be clearer in. Which half is the safety net and which is the
// lesson depends on who is writing: the pair is (English + X) either way, and an
// English speaker learning French wants the French half for the same reason a
// Mandarin speaker learning English wants the English one. Petal cannot tell
// them apart from a chat message, and doesn't need to — every other explanation
// surface already gives both (the card's English body, the seeded bubble in the
// pair language). The answer that goes deepest into the "why" was the one place
// that didn't.
//
// The blank line between the halves is a contract with the client: AskPetal.tsx
// splits on the first one to render her language prominently and the English
// beneath it, mirroring the companion's bubble. A model that ignores the
// instruction and writes one language degrades to a single plain block — the
// answer is still readable, which is why the split is a rendering nicety and
// never a parse the reply depends on.
const askPetalSystemTemplate = `You are Petal, a warm and patient English writing tutor helping someone who is learning English ` +
`as a second language. You are currently discussing a specific writing suggestion.
@@ -154,15 +173,21 @@ Suggestion context:
- Initial explanation: "%[4]s"
- Surrounding paragraph: "%[5]s"
The user wants to understand this suggestion better. Detect the language of the user's message ` +
`and respond in that same language. If they write in %[6]s, respond entirely in ` +
`%[6]s. If they write in English, respond in English. Never mix languages in a single response.
The user wants to understand this suggestion better. Answer in BOTH languages, every time, ` +
`whichever language they asked their question in: first the whole answer in %[6]s, then the ` +
`same answer again in English. Separate the two with a single blank line. Do not label them, ` +
`do not use a blank line anywhere else, and do not mix the two languages within one half — ` +
`each half is complete on its own.
One of those two languages is the one they are surest in and the other is the one they are ` +
`working in — you do not know which way round, so give both and let them choose. Both halves ` +
`say the same thing: do not put a point in one that is missing from the other.
Explain clearly and kindly. Use simple language appropriate to the user's message. Give examples ` +
`when helpful. If they ask "why" (or "%[7]s"), explain the grammar rule or idiom behind it. ` +
`If they suggest an alternative phrasing, evaluate it honestly.
Keep responses concise (2-4 sentences). This is a chat, not an essay. Be encouraging — ` +
Keep each half concise (2-3 sentences). This is a chat, not an essay. Be encouraging — ` +
`learning a language is hard and they're doing great.`
// AskPetalSystemPrompt fills the tutor prompt with one suggestion's context and
+145
View File
@@ -0,0 +1,145 @@
package suggestions
import (
"crypto/sha256"
"encoding/hex"
"strings"
"unicode"
)
// Chunking splits a document into sentence-sized units so a re-check can ask the
// model only about the sentences that actually changed. Accepting one edit used
// to re-run the whole document: every card vanished, came back with a new id and
// a freshly-worded explanation, and spans re-merged into different shapes. The
// sentences she didn't touch have nothing new to say about themselves, so their
// suggestions are simply kept (see reconcilePending).
//
// A chunk's identity is its hash, not its position — she inserts a paragraph at
// the top and every sentence below keeps its suggestions.
// chunk is one sentence of the document, with the hash that identifies it.
type chunk struct {
text string
hash string
}
// asciiTerminators end a sentence only when whitespace (or the end of the text)
// follows, so "3.5" and "Ms." don't split mid-word — a wrong split costs only a
// slightly smaller chunk, but a split inside a number would churn its hash on
// every keystroke around it.
const asciiTerminators = ".!?"
// cjkTerminators end a sentence outright: Chinese runs sentences together with
// no space after 。, and she writes in both languages in one document.
const cjkTerminators = "。!?"
// closers are swallowed into the sentence they close, so the quote mark travels
// with the sentence rather than opening the next one.
const closers = `)]}"'’”」』`
// splitChunks divides text into sentences, dropping whitespace-only runs.
// Newlines always break a chunk, so a list or a line of dialogue is its own unit.
//
// `salt` distinguishes two *readings* of the same sentence. The grammar
// checkpoint's advice depends on the document's tone — the same line gets
// different notes as an academic essay than as a journal entry — so switching
// tone must re-open every sentence rather than serve back advice written for the
// old register.
func splitChunks(text, salt string) []chunk {
var out []chunk
runes := []rune(text)
start := 0
add := func(end int) {
if s := string(runes[start:end]); strings.TrimSpace(s) != "" {
out = append(out, chunk{text: s, hash: hashChunk(s, salt)})
}
start = end
}
for i := 0; i < len(runes); i++ {
r := runes[i]
if r == '\n' {
add(i + 1)
continue
}
cjk := strings.ContainsRune(cjkTerminators, r)
if !cjk && !strings.ContainsRune(asciiTerminators, r) {
continue
}
// Swallow a run of terminators ("?!", "…") and any closing punctuation.
j := i + 1
for j < len(runes) && (strings.ContainsRune(asciiTerminators+cjkTerminators+closers, runes[j])) {
j++
}
if cjk || j >= len(runes) || unicode.IsSpace(runes[j]) {
add(j)
i = j - 1
}
}
if start < len(runes) {
add(len(runes))
}
return out
}
// hashChunk identifies a sentence by its content under the same normalization
// the suppression logic uses: quote style and whitespace runs churn constantly
// (the editor rewrites quotes as she types, a paragraph reflows) and none of
// that changes what the sentence says, so none of it should cost a re-check.
func hashChunk(s, salt string) string {
sum := sha256.Sum256([]byte(salt + "\x00" + normalizeForDedup(s)))
return hex.EncodeToString(sum[:])[:16]
}
// hashSet indexes chunks by hash — "is this sentence in the document?"
func hashSet(chunks []chunk) map[string]bool {
out := make(map[string]bool, len(chunks))
for _, c := range chunks {
out[c.hash] = true
}
return out
}
// changedChunks returns the chunks whose hash wasn't in the last checked set,
// in document order and deduplicated — a sentence repeated verbatim is one
// question, not two.
func changedChunks(chunks []chunk, checked map[string]bool) []chunk {
seen := make(map[string]bool, len(chunks))
var out []chunk
for _, c := range chunks {
if checked[c.hash] || seen[c.hash] {
continue
}
seen[c.hash] = true
out = append(out, c)
}
return out
}
// joinChunks renders a chunk set as the text to hand the model: one sentence per
// line, so two sentences pulled from opposite ends of the document don't read as
// one run-on.
func joinChunks(chunks []chunk) string {
parts := make([]string, 0, len(chunks))
for _, c := range chunks {
parts = append(parts, strings.TrimSpace(c.text))
}
return strings.Join(parts, "\n")
}
// chunkFor names the sentence a suggestion belongs to: the first chunk whose
// text contains the flagged span. Returns "" when the span straddles a sentence
// boundary or the model paraphrased what it quoted — such a row is re-examined
// on every pass rather than cached, which is the safe direction.
func chunkFor(original string, chunks []chunk) string {
o := normalizeForDedup(original)
if o == "" {
return ""
}
for _, c := range chunks {
if strings.Contains(normalizeForDedup(c.text), o) {
return c.hash
}
}
return ""
}
+118
View File
@@ -0,0 +1,118 @@
package suggestions
import (
"strings"
"testing"
)
func texts(chunks []chunk) []string {
out := make([]string, 0, len(chunks))
for _, c := range chunks {
out = append(out, strings.TrimSpace(c.text))
}
return out
}
func TestSplitChunks(t *testing.T) {
cases := []struct {
name string
in string
want []string
}{
{
name: "plain sentences",
in: "I has two apple. She go to market yesterday! Why?",
want: []string{"I has two apple.", "She go to market yesterday!", "Why?"},
},
{
// A decimal must not split, or the sentence's identity would churn
// while she types the number.
name: "decimals stay whole",
in: "It costs 3.50 today. Tomorrow, more.",
want: []string{"It costs 3.50 today.", "Tomorrow, more."},
},
{
name: "closing quote travels with its sentence",
in: `He said "early," and left. She stayed.`,
want: []string{`He said "early," and left.`, "She stayed."},
},
{
// Chinese runs sentences together with no space after 。 — she writes
// in both languages in one document.
name: "cjk terminators split without a space",
in: "我想说这句话。但是不知道用英语怎么说。",
want: []string{"我想说这句话。", "但是不知道用英语怎么说。"},
},
{
name: "newlines break chunks",
in: "A list item\nAnother item\n",
want: []string{"A list item", "Another item"},
},
{
name: "blank runs are dropped",
in: "\n\n \nOnly this.\n\n",
want: []string{"Only this."},
},
{
name: "trailing fragment is its own chunk",
in: "Done. Still writing",
want: []string{"Done.", "Still writing"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := texts(splitChunks(tc.in, ""))
if len(got) != len(tc.want) {
t.Fatalf("want %q, got %q", tc.want, got)
}
for i := range got {
if got[i] != tc.want[i] {
t.Fatalf("chunk %d: want %q, got %q", i, tc.want[i], got[i])
}
}
})
}
}
// A sentence's identity survives the churn that doesn't change what it says:
// the editor rewrites quotes as she types, and a paragraph reflows.
func TestChunkIdentityIgnoresCosmeticChurn(t *testing.T) {
a := splitChunks(`She said "hello" softly.`, "")
b := splitChunks("She said “hello” softly.", "")
if len(a) != 1 || len(b) != 1 {
t.Fatalf("want one chunk each, got %d and %d", len(a), len(b))
}
if a[0].hash != b[0].hash {
t.Fatalf("quote/whitespace churn changed the sentence's identity")
}
if same := splitChunks(`She said "hello" softly.`, "academic"); same[0].hash == a[0].hash {
t.Fatalf("a different tone must be a different reading of the sentence")
}
}
func TestChangedChunksAndLookup(t *testing.T) {
chunks := splitChunks("One thing. Another thing. One thing.", "")
if len(chunks) != 3 {
t.Fatalf("want 3 chunks, got %d", len(chunks))
}
// A repeated sentence is one question, not two.
if got := changedChunks(chunks, nil); len(got) != 2 {
t.Fatalf("want 2 distinct changed chunks, got %d", len(got))
}
checked := hashSet(chunks[:1])
changed := changedChunks(chunks, checked)
if len(changed) != 1 || strings.TrimSpace(changed[0].text) != "Another thing." {
t.Fatalf("want only the unread sentence, got %q", texts(changed))
}
if chunkFor("Another", chunks) != chunks[1].hash {
t.Fatalf("span was attributed to the wrong sentence")
}
// A span the document doesn't contain has no sentence, so it is never cached.
if chunkFor("nowhere in here", chunks) != "" {
t.Fatalf("unanchorable span should have no chunk")
}
}
+202 -77
View File
@@ -56,6 +56,7 @@ func (h *Handler) RegisterDocRoutes(r chi.Router) {
r.Post("/{id}/collocation", h.collocation)
r.Post("/{id}/rewrite", h.rewrite)
r.Get("/{id}/suggestions", h.listForDoc)
r.Get("/{id}/settled", h.listSettled)
}
// Routes returns the router mounted at /api/suggestions for per-suggestion
@@ -160,16 +161,21 @@ func (h *Handler) mechanics(w http.ResponseWriter, r *http.Request) {
httputil.WriteJSON(w, http.StatusOK, out)
}
// replaceMechanics swaps the document's pending offline rows for the supplied
// findings in one transaction, leaving the LLM families and actioned rows
// untouched. Findings the user already accepted or dismissed are suppressed (the
// detector has no memory between runs), and malformed spans are skipped.
// replaceMechanics brings the document's pending offline rows in line with the
// supplied findings in one transaction, leaving the LLM families and actioned
// rows untouched. Findings the user already accepted or dismissed are suppressed
// (the detector has no memory between runs), and malformed spans are skipped.
//
// The DELETE is scoped by *source*, not by type: the rule pack owns both the
// mechanics family and its share of the collocation family, and every run is a
// full recompute of the document, so everything it wrote last time goes. Scoping
// by type instead would strand offline collocations the current text no longer
// warrants — the one row nobody would ever replace.
// A finding the detector still reports keeps its existing row — same id, same
// created_at — and only its offsets move. This pass fires 250 ms after a
// keystroke, so deleting and re-inserting the family would hand every card a new
// identity several times a sentence: the rail would remount, a card expanded for
// Ask Petal would collapse under her, and the arrival chime would re-fire.
//
// The scope is *source*, not type: the rule pack owns both the mechanics family
// and its share of the collocation family, and every run is a full recompute of
// the document. Scoping by type instead would strand offline collocations the
// current text no longer warrants — the one row nobody would ever replace.
func (h *Handler) replaceMechanics(docID string, findings []mechanicsFinding) error {
tx, err := h.DB.Begin()
if err != nil {
@@ -177,18 +183,18 @@ func (h *Handler) replaceMechanics(docID string, findings []mechanicsFinding) er
}
defer tx.Rollback()
if _, err := tx.Exec(
`DELETE FROM suggestions WHERE doc_id = ? AND status = ? AND source = ?`,
docID, db.SuggestionStatusPending, db.SuggestionSourceLocal,
); err != nil {
existing, err := loadPending(tx, docID, "source = '"+db.SuggestionSourceLocal+"'")
if err != nil {
return err
}
index := indexByEdit(existing)
sup, err := buildSuppressor(tx, docID)
if err != nil {
return err
}
kept := make(map[string]bool, len(existing))
for _, f := range findings {
if f.From < 0 || f.To <= f.From || strings.TrimSpace(f.Original) == "" {
continue // malformed span — the client re-anchors by string anyway
@@ -196,16 +202,34 @@ func (h *Handler) replaceMechanics(docID string, findings []mechanicsFinding) er
if sup.suppressed(f.Original, f.Replacement) {
continue
}
typ := localType(f.Type)
if row, ok := index.take(f.Original, f.Replacement, f.From); ok {
kept[row.id] = true
if err := reposition(tx, row, f.From, f.To, ""); err != nil {
return err
}
continue
}
if _, err := tx.Exec(
`INSERT INTO suggestions (doc_id, from_pos, to_pos, original, replacement, explanation, type, source)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)`,
docID, f.From, f.To, f.Original, f.Replacement, f.Explanation,
localType(f.Type), db.SuggestionSourceLocal,
typ, db.SuggestionSourceLocal,
); err != nil {
return err
}
}
// Whatever the detector no longer reports, she has fixed.
for _, row := range existing {
if kept[row.id] {
continue
}
if _, err := tx.Exec(`DELETE FROM suggestions WHERE id = ?`, row.id); err != nil {
return err
}
}
return tx.Commit()
}
@@ -258,11 +282,64 @@ func (h *Handler) runPass(w http.ResponseWriter, r *http.Request, limiter *llm.R
return
}
// Nothing to analyze on an empty document — skip the LLM round-trip.
// Nothing to analyze on an empty document — skip the LLM round-trip. The
// family's rows go with the text they were about.
if strings.TrimSpace(contentText) == "" {
httputil.WriteJSON(w, http.StatusOK, []db.Suggestion{})
if err := h.reconcilePending(docID, contentText, pairLang, nil, scope, nil, nil, false); err != nil {
httputil.ServerError(w, err)
return
}
out, err := h.fetchPending(userID, docID)
if err != nil {
httputil.ServerError(w, err)
return
}
httputil.WriteJSON(w, http.StatusOK, out)
return
}
// Decide what to ask about before spending anything: a chunked pass asks only
// about the sentences that changed since it last read the document, and when
// none did it doesn't call the model at all — nor consume its rate-limit slot,
// so the next real edit isn't throttled by a check that had nothing to do.
//
// Only a chunked pass consults that record, so only it needs the tone folded
// into a sentence's identity.
salt := ""
if scope.chunked {
salt = tone
}
chunks := splitChunks(contentText, salt)
askText, fresh := contentText, chunks
if scope.chunked {
checked, err := h.checkedChunks(docID, scope.family)
if err != nil {
httputil.ServerError(w, err)
return
}
changed := changedChunks(chunks, checked)
if len(changed) == 0 {
// Every sentence has already been read. Drop the rows whose sentence is
// gone, keep the rest exactly as they are, and answer immediately.
if err := h.reconcilePending(docID, contentText, pairLang, nil, scope, chunks, nil, false); err != nil {
httputil.ServerError(w, err)
return
}
out, err := h.fetchPending(userID, docID)
if err != nil {
httputil.ServerError(w, err)
return
}
httputil.WriteJSON(w, http.StatusOK, out)
return
}
// When every sentence is new — a first pass, a paste, a tone switch — hand
// over the document verbatim so the model reads it with its paragraphing
// intact. Otherwise send just the delta, one sentence per line.
if len(changed) < len(hashSet(chunks)) {
askText, fresh = joinChunks(changed), changed
}
}
ok, _, slotAt := limiter.Allow(docID)
if !ok {
@@ -277,7 +354,7 @@ func (h *Handler) runPass(w http.ResponseWriter, r *http.Request, limiter *llm.R
return
}
raw, err := run(r.Context(), h.Client, contentText, tone, llm.LangFor(pairLang))
raw, err := run(r.Context(), h.Client, askText, tone, llm.LangFor(pairLang))
if err != nil {
// Allow ran before the model call, so a failed pass would otherwise hold
// the per-document slot for the full interval — stranding the frontend's
@@ -287,7 +364,9 @@ func (h *Handler) runPass(w http.ResponseWriter, r *http.Request, limiter *llm.R
return
}
if err := h.replacePending(docID, contentText, raw, scope); err != nil {
// A whole-document pass re-read everything, so every one of its rows is up for
// re-proposal; a chunked pass only puts the sentences it asked about in play.
if err := h.reconcilePending(docID, contentText, pairLang, raw, scope, chunks, fresh, !scope.chunked); err != nil {
httputil.ServerError(w, err)
return
}
@@ -308,78 +387,49 @@ func (h *Handler) runPass(w http.ResponseWriter, r *http.Request, limiter *llm.R
// inserts. The grammar checkpoint and voice pass each own a disjoint family, so
// running one never disturbs the other's pending flags.
type pendingScope struct {
deleteWhere string // extra WHERE clause scoping the DELETE to this family
deleteWhere string // extra WHERE clause scoping this pass to its own family
forceType string // if set, every inserted row gets this type; else normalizeType
// family keys the sentences this pass has already read (see checked_chunks).
family string
// chunked passes re-read only the sentences that changed. True for the typing-
// cadence grammar checkpoint, which fires constantly and must feel still;
// false for the explicit whole-document passes, where she pressed a button
// asking for a fresh read of everything.
chunked bool
}
// Every scope below is confined to source='llm'. The offline rule pack replaces
// its own rows wholesale on each edit (see replaceMechanics) and its findings
// Every scope below is confined to source='llm'. The offline rule pack owns its
// own rows and recomputes them on each edit (see replaceMechanics); its findings
// must survive all three model passes — including the collocation coach, which
// now shares the collocation family with it.
var (
// grammarScope owns the grammar/phrasing/idiom/clarity flags — everything but
// the other self-owned families (voice, collocation), which run on their own
// cadence/pass and must survive a grammar checkpoint. Notably the offline pass
// writes its rows in the same /check request just before this DELETE runs, so
// the source clause is also what keeps them alive.
grammarScope = pendingScope{deleteWhere: "source = 'llm' AND type NOT IN ('voice','collocation')", forceType: ""}
// voiceScope owns the model's voice flags only.
voiceScope = pendingScope{deleteWhere: "source = 'llm' AND type = 'voice'", forceType: db.SuggestionTypeVoice}
// writes its rows in the same /check request just before this pass reconciles,
// so the source clause is also what keeps them alive.
grammarScope = pendingScope{
deleteWhere: "source = 'llm' AND type NOT IN ('voice','collocation')",
family: "grammar",
chunked: true,
}
// voiceScope owns the model's voice flags only. Voice is a property of the
// document as a whole — a sentence isn't inconsistent with itself — so this
// pass always reads everything.
voiceScope = pendingScope{
deleteWhere: "source = 'llm' AND type = 'voice'",
forceType: db.SuggestionTypeVoice,
family: "voice",
}
// collocationScope owns the model's collocation flags only — the rule pack's
// share of the same family is left standing.
collocationScope = pendingScope{deleteWhere: "source = 'llm' AND type = 'collocation'", forceType: db.SuggestionTypeCollocation}
collocationScope = pendingScope{
deleteWhere: "source = 'llm' AND type = 'collocation'",
forceType: db.SuggestionTypeCollocation,
family: "collocation",
}
)
// replacePending swaps a document's pending suggestions within one family for a
// fresh batch in a single transaction. Accepted/rejected suggestions and the
// other family's pending rows are left untouched.
//
// Suggestions touching a sentence the user already settled are suppressed from
// the fresh batch (see suppressor): not just the identical edit re-proposed, but
// reversals and re-polishing of the model's own just-accepted output — the
// "fickle, keeps going back and forth on a few sentences" behavior. The model has
// no memory between passes, so without this it re-opens resolved sentences every
// checkpoint.
func (h *Handler) replacePending(docID, contentText string, raw []llm.RawSuggestion, scope pendingScope) error {
tx, err := h.DB.Begin()
if err != nil {
return err
}
defer tx.Rollback()
if _, err := tx.Exec(
`DELETE FROM suggestions WHERE doc_id = ? AND status = ? AND `+scope.deleteWhere,
docID, db.SuggestionStatusPending,
); err != nil {
return err
}
sup, err := buildSuppressor(tx, docID)
if err != nil {
return err
}
for _, s := range raw {
if sup.suppressed(s.Original, s.Replacement) {
continue
}
typ := scope.forceType
if typ == "" {
typ = normalizeType(s.Type)
}
from, to := locate(contentText, s.Original)
if _, err := tx.Exec(
`INSERT INTO suggestions (doc_id, from_pos, to_pos, original, replacement, explanation, type, source)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)`,
docID, from, to, s.Original, s.Replacement, s.Explanation, typ, db.SuggestionSourceLLM,
); err != nil {
return err
}
}
return tx.Commit()
}
// dedupQuoteReplacer folds every straight/curly single- and double-quote variant
// (and backtick/acute accent) onto one canonical character. The editor and the
// model both rewrite quotes between passes — a sentence accepted with "…" comes
@@ -518,6 +568,76 @@ func (h *Handler) listForDoc(w http.ResponseWriter, r *http.Request) {
httputil.WriteJSON(w, http.StatusOK, out)
}
// listSettled returns the normalized originals of every edit the user has
// already accepted or dismissed on this document — the same spans buildSuppressor
// drops on the server, handed to the client so its instant rule-pack pass can
// drop them too.
//
// Without this the offline half of the loop has no memory. The rule pack detects
// from the text alone and re-runs 250 ms after a keystroke, so a dismissed "the
// the" comes straight back the moment she types anywhere in the document; the
// server's reply then removes it again. That flicker is the visible symptom, but
// the real one is worse: with the server unreachable — the case the rule pack
// exists for — the reply never comes and a card she dismissed simply stays.
//
// Only the originals are sent. Replacements are the model's words, not hers, and
// the client only needs to answer "has she settled this span?"
func (h *Handler) listSettled(w http.ResponseWriter, r *http.Request) {
out, err := h.fetchSettled(auth.UserID(r.Context()), chi.URLParam(r, "id"))
if err != nil {
httputil.ServerError(w, err)
return
}
httputil.WriteJSON(w, http.StatusOK, settledResponse{Originals: out})
}
// settledResponse wraps the list so the endpoint can grow a second field without
// breaking a client that reads a bare array.
type settledResponse struct {
Originals []string `json:"originals"`
}
// fetchSettled loads the distinct normalized originals of the document's actioned
// rows. Scoped through documents for the same reason fetchPending is: an original
// is a quotation of her writing.
func (h *Handler) fetchSettled(userID, docID string) ([]string, error) {
rows, err := h.DB.Query(
`SELECT DISTINCT s.original
FROM suggestions s
JOIN documents d ON d.id = s.doc_id
WHERE s.doc_id = ? AND d.user_id = ? AND s.status IN (?, ?)`,
docID, userID, db.SuggestionStatusAccepted, db.SuggestionStatusRejected,
)
if err != nil {
return nil, err
}
defer rows.Close()
// DISTINCT is on the raw text; normalizing can collapse two rows into one, so
// dedupe again on this side to keep the payload honest.
seen := map[string]struct{}{}
out := []string{}
for rows.Next() {
var original string
if err := rows.Scan(&original); err != nil {
return nil, err
}
norm := normalizeForDedup(original)
if norm == "" {
continue
}
if _, dup := seen[norm]; dup {
continue
}
seen[norm] = struct{}{}
out = append(out, norm)
}
if err := rows.Err(); err != nil {
return nil, err
}
return out, nil
}
// fetchPending loads a document's pending suggestions, joined through documents
// so the rows are only reachable by the document's owner. A suggestion quotes the
// sentence it corrects, so an unscoped read here would leak document text to
@@ -707,6 +827,11 @@ func locate(contentText, original string) (int, int) {
// normalizeType maps the model's type string onto a valid suggestion type,
// defaulting unknown values to grammar so a stray label never trips the CHECK.
//
// 'translate' is absent on purpose, and stays absent even though the type now
// exists: it is decided from the span (see language.go), never taken from the
// model. A model that volunteers the label anyway lands on grammar here and is
// then promoted — or not — on the evidence.
func normalizeType(t string) string {
switch strings.ToLower(strings.TrimSpace(t)) {
case db.SuggestionTypeGrammar, db.SuggestionTypePhrasing, db.SuggestionTypeIdiom, db.SuggestionTypeClarity, db.SuggestionTypeCollocation:
+27 -2
View File
@@ -20,10 +20,17 @@ import (
type stubClient struct {
response string
calls int
// The full prompt of the most recent call, so a test can assert which
// sentences a chunked pass actually asked about.
lastPrompt string
}
func (s *stubClient) Complete(_ context.Context, _ llm.CompletionRequest) (string, error) {
func (s *stubClient) Complete(_ context.Context, req llm.CompletionRequest) (string, error) {
s.calls++
s.lastPrompt = ""
for _, m := range req.Messages {
s.lastPrompt += m.Content + "\n"
}
return s.response, nil
}
@@ -64,6 +71,17 @@ func newTestServer(t *testing.T, client llm.LLMClient) (http.Handler, string, *H
return authed, docID, h
}
// setDocText rewrites the seeded document, standing in for the writer editing.
// The grammar checkpoint only asks the model about sentences that changed since
// it last read the document, so a test that wants a second real pass has to
// change something first — as she always has.
func setDocText(t *testing.T, h *Handler, docID, text string) {
t.Helper()
if _, err := h.DB.Exec(`UPDATE documents SET content_text = ? WHERE id = ?`, text, docID); err != nil {
t.Fatalf("update doc text: %v", err)
}
}
func do(t *testing.T, srv http.Handler, method, path, body string) *httptest.ResponseRecorder {
t.Helper()
var r *http.Request
@@ -183,6 +201,7 @@ func TestFickleEditsSuppressed(t *testing.T) {
]}`}
srv, docID, h := newTestServer(t, client)
h.Limit = llm.NewRateLimiter(0)
setDocText(t, h, docID, `He left "early," because of the rain. The cat always have a calm face.`)
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var got []db.Suggestion
@@ -194,6 +213,10 @@ func TestFickleEditsSuppressed(t *testing.T) {
do(t, srv, http.MethodPost, "/suggestions/"+s.ID+"/accept", "")
}
// Both edits are now in the document, which is what re-opens those sentences
// for a second reading.
setDocText(t, h, docID, `He left "early," due to the rain. The cat always has a calm face.`)
// Reversal of the first accept (note the " → ' quote churn) and a re-polish of
// the second accept must both be dropped; only the unrelated edit survives.
client.response = `{"suggestions":[
@@ -314,7 +337,9 @@ func TestCollocationPassCoexists(t *testing.T) {
t.Fatalf("collocation response should carry all three families, got %+v", got)
}
// A grammar checkpoint must NOT wipe the voice or collocation flags.
// A grammar checkpoint must NOT wipe the voice or collocation flags. She fixes
// the flagged sentence, so its own grammar row goes and nothing replaces it.
setDocText(t, h, docID, "I have two apples.")
client.response = `{"suggestions":[]}`
rec = do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
if rec.Code != http.StatusOK {
+10
View File
@@ -116,4 +116,14 @@ func TestSuggestionIsolation(t *testing.T) {
if rec.Code != http.StatusNoContent {
t.Fatalf("owner accept = %d, want 204 (body: %s)", rec.Code, rec.Body)
}
// That accept created a settled span, which is the other read of this table.
// It carries originals only — but an original is a verbatim quotation of her
// sentence, so it is the same leak as the pending list through a smaller hole.
if got := getSettled(t, owner, docID); len(got) != 1 {
t.Fatalf("owner should see their own settled span, got %v", got)
}
if got := getSettled(t, stranger, docID); len(got) != 0 {
t.Fatalf("stranger read %d settled span(s) (leaking %q)", len(got), got[0])
}
}
+174
View File
@@ -0,0 +1,174 @@
package suggestions
import (
"strings"
"unicode"
)
// Telling her language from English, well enough to label a card.
//
// When the checkpoint quotes a span she wrote in her own language and hands back
// an English rendering, that is not a correction — nothing was wrong with what
// she wrote — and it should not be filed under 'clarity'. The label is derived
// here rather than asked of the model: a type is structural, and a model that
// re-reasons every pass would drift between labels for the same sentence.
//
// The failure mode is deliberately cheap. Getting this wrong changes a card's
// coloured pill and nothing else — the replacement, the explanation and the
// Accept button are identical either way — so a heuristic is the right tool. It
// is written to under-claim: a span it isn't sure about stays whatever the model
// called it.
//
// The two pair families need genuinely different tests, and pretending otherwise
// would be the bug:
//
// - zh is a different script. Counting Han runes is close to certain.
// - pt-PT, fr and es share the Latin alphabet with English, where no such
// signal exists. Those fall back to function words — the short, extremely
// common words a sentence in that language can hardly avoid and an English
// sentence has no reason to contain.
// isTranslation reports whether this edit is her own language rendered into
// English, rather than a correction to her English. Both halves must hold: the
// quoted span reads as the pair language, and what Petal offers back reads as
// English. The second half matters — a Chinese span rewritten into different
// Chinese is something else entirely, and Petal has no business calling it a
// translation.
func isTranslation(original, replacement, pairLang string) bool {
if strings.TrimSpace(original) == "" || strings.TrimSpace(replacement) == "" {
return false
}
return readsAsPairLang(original, pairLang) && readsAsEnglish(replacement)
}
// readsAsPairLang reports whether s is predominantly in the writer's language.
func readsAsPairLang(s, pairLang string) bool {
switch normalizePairLang(pairLang) {
case "zh":
han, latin := scriptCounts(s)
// Predominantly, not merely partly: one Chinese word inside an English
// sentence is a vocabulary question, and the sentence around it is still
// English prose with its own grammar to correct. Two runes is the floor
// because a single Han character is as likely to be a stray keystroke.
return han >= 2 && han > latin
case "pt-PT", "fr", "es":
return distinctMarkers(s, latinMarkers[normalizePairLang(pairLang)]) >= 2
}
// A pair Petal has no test for. Say no: an unlabelled card is a card that
// reads as it did yesterday, and a wrongly-labelled one is a new defect.
return false
}
// readsAsEnglish reports whether s is English prose rather than more of her own
// language. It is not a language identifier — it only has to separate "English"
// from "the pair language", and it is only ever asked about text Petal itself
// generated, so the bar is low on purpose: Latin letters present, and not
// swamped by another script.
func readsAsEnglish(s string) bool {
han, latin := scriptCounts(s)
return latin > 0 && latin > han
}
// normalizePairLang folds the stored `users.pair_lang` into the codes below.
// Empty (a document whose owner has no pair recorded) falls through to no test.
func normalizePairLang(pairLang string) string {
switch p := strings.ToLower(strings.TrimSpace(pairLang)); p {
case "zh", "zh-cn", "zh-hans":
return "zh"
case "pt", "pt-pt":
return "pt-PT"
case "fr", "fr-fr":
return "fr"
case "es", "es-es":
return "es"
default:
return p
}
}
// scriptCounts counts Han runes and ASCII letters. Everything else — digits,
// punctuation, spaces, emoji — is ignored, so trailing 。or a stray comma
// changes nothing.
func scriptCounts(s string) (han, latin int) {
for _, r := range s {
switch {
case unicode.Is(unicode.Han, r):
han++
case r < unicode.MaxASCII && unicode.IsLetter(r):
latin++
}
}
return han, latin
}
// distinctMarkers counts how many *different* marker words appear in s. Distinct
// rather than total: "que ... que" is one writer's habit, while "eu quero" is two
// independent pieces of evidence.
func distinctMarkers(s string, markers map[string]bool) int {
if len(markers) == 0 {
return 0
}
seen := map[string]bool{}
for _, w := range strings.FieldsFunc(strings.ToLower(s), func(r rune) bool {
// Split on anything that isn't a letter, so punctuation and digits are
// separators. Apostrophes included: French elision (j'ai, n'est) should
// yield its parts.
return !unicode.IsLetter(r)
}) {
if markers[w] {
seen[w] = true
}
}
return len(seen)
}
// Function words that a sentence in each Latin pair can hardly avoid.
//
// Curated against English, not for coverage: every entry here is a word an
// English sentence has essentially no reason to contain, which is why the lists
// omit plenty of far more common words. Deliberately absent — each of them a
// false positive waiting to happen — is anything that is *also* an English word:
// the pan-Romance shorts (a, o, e, as, no, on, en, de, se, na, mi, son, era,
// plus, pour, si, ma, ce, ne), Portuguese "do", Spanish "con", "ya" and "todo".
// Dropping "con" costs the Spanish list one of its commonest words, and that is
// the right trade — a marker that fires on English corroborates the wrong
// answer, which is worse than a sentence Petal declines to label.
//
// A single marker is not enough (see readsAsPairLang), so these lists are read
// as evidence to be corroborated rather than as a decision.
var latinMarkers = map[string]map[string]bool{
"fr": words(
"je", "tu", "il", "elle", "ils", "elles", "nous", "vous", "est", "sont",
"était", "étais", "une", "des", "les", "du", "dans", "avec", "que", "qui",
"mais", "très", "être", "avoir", "pas", "cette", "cet", "ces", "mon",
"mes", "notre", "votre", "leur", "aussi", "alors", "parce", "comme",
"beaucoup", "toujours", "jamais", "quand", "bien", "chose", "temps",
"moi", "toi", "lui", "peux", "veux", "sais", "faire", "dit", "aujourd",
"hui", "quelque", "chez", "tout", "tous", "rien", "déjà", "encore",
),
"pt-PT": words(
"eu", "você", "ele", "ela", "eles", "elas", "nós", "são", "uma", "os",
"da", "dos", "das", "com", "que", "mas", "muito", "não", "meu",
"minha", "seu", "sua", "isso", "este", "esta", "está", "estou", "quero",
"também", "quando", "porque", "coisa", "tempo", "fazer", "sempre",
"nunca", "bem", "obrigado", "obrigada", "gosto", "tenho", "tem", "foi",
"ser", "ter", "mais", "já", "ainda", "aqui", "ali", "nada", "tudo",
"todos", "para", "pela", "pelo", "sobre", "assim",
),
"es": words(
"yo", "él", "ella", "ellos", "ellas", "nosotros", "una", "los", "las",
"del", "que", "pero", "muy", "esto", "esta", "este", "está",
"estoy", "quiero", "también", "cuando", "porque", "cosa", "tiempo",
"hacer", "siempre", "nunca", "bien", "gracias", "tengo", "tiene", "fue",
"ser", "tener", "más", "aquí", "allí", "nada", "todos",
"para", "sobre", "así", "hola", "señor", "usted", "muchas",
),
}
func words(list ...string) map[string]bool {
out := make(map[string]bool, len(list))
for _, w := range list {
out[w] = true
}
return out
}
+176
View File
@@ -0,0 +1,176 @@
package suggestions
import "testing"
// The flagship case, and the ones next to it that must NOT become translations.
func TestIsTranslation(t *testing.T) {
cases := []struct {
name string
original string
replacement string
pairLang string
want bool
}{
{
// The sentence from the UX review, verbatim.
name: "whole Chinese sentence rendered into English",
original: "我想说这句话但是不知道用英语怎么说。",
replacement: "I want to say this but I don't know how to say it in English.",
pairLang: "zh",
want: true,
},
{
name: "ordinary English correction is not a translation",
original: "She goes to market yesterday",
replacement: "She went to the market yesterday",
pairLang: "zh",
want: false,
},
{
// One Chinese word inside English prose. The sentence around it is
// still English with its own grammar to fix, and calling the card a
// translation would mislabel a grammar fix.
name: "single Chinese word inside an English sentence",
original: "I bought a 苹果 at the store",
replacement: "I bought an apple at the store",
pairLang: "zh",
want: false,
},
{
name: "a lone stray Han rune is not a sentence",
original: "的",
replacement: "of",
pairLang: "zh",
want: false,
},
{
// Chinese in, Chinese out: whatever this is, Petal is not translating.
name: "Chinese rewritten as Chinese",
original: "我想说这句话",
replacement: "我要说这句话",
pairLang: "zh",
want: false,
},
{
// The same Chinese span, but the writer is on the French pair. Petal
// has no business offering to translate a language she never claimed.
name: "Chinese span on a non-zh pair",
original: "我想说这句话但是不知道用英语怎么说。",
replacement: "I want to say this in English.",
pairLang: "fr",
want: false,
},
{
name: "French sentence rendered into English",
original: "Je ne sais pas comment le dire en anglais.",
replacement: "I don't know how to say it in English.",
pairLang: "fr",
want: true,
},
{
name: "Portuguese sentence rendered into English",
original: "Eu quero dizer isso mas não sei como.",
replacement: "I want to say this but I don't know how.",
pairLang: "pt-PT",
want: true,
},
{
name: "Spanish sentence rendered into English",
original: "Yo quiero decir esto pero no sé cómo.",
replacement: "I want to say this but I don't know how.",
pairLang: "es",
want: true,
},
{
// A single marker is not evidence. "Que" appears in English writing
// about other languages, in names, in quoted phrases.
name: "one Latin marker is not enough",
original: "The word que confused me",
replacement: "The word que confuses me",
pairLang: "pt-PT",
want: false,
},
{
// The words most likely to sink this heuristic: English function words
// that are also Romance function words. They are kept out of the lists
// precisely so this sentence stays a grammar fix.
name: "English full of pan-Romance lookalikes",
original: "I do not know if a con man on the plus side as no era",
replacement: "I do not know whether a con man, on the plus side, is no era",
pairLang: "es",
want: false,
},
{
name: "English with a borrowed French phrase stays English",
original: "It was a pas de deux, more or less",
replacement: "It was a pas de deux, more or less.",
pairLang: "fr",
want: false,
},
{
name: "empty replacement (an awareness-only finding)",
original: "我想说这句话但是不知道用英语怎么说。",
replacement: "",
pairLang: "zh",
want: false,
},
{
// A document whose owner has no pair recorded. No test, no label.
name: "no pair language",
original: "我想说这句话但是不知道用英语怎么说。",
replacement: "I want to say this in English.",
pairLang: "",
want: false,
},
{
// An unshipped pair. Same rule: decline rather than guess.
name: "unknown pair language",
original: "Ich weiß nicht wie man das sagt.",
replacement: "I don't know how to say that.",
pairLang: "de",
want: false,
},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := isTranslation(c.original, c.replacement, c.pairLang); got != c.want {
t.Errorf("isTranslation(%q, %q, %q) = %v, want %v",
c.original, c.replacement, c.pairLang, got, c.want)
}
})
}
}
// pair_lang is stored as the pack code, but a stored value has drifted before
// (see the picker's history), so the fold is tested rather than assumed.
func TestNormalizePairLang(t *testing.T) {
for in, want := range map[string]string{
"zh": "zh", "zh-CN": "zh", "ZH": "zh",
"pt": "pt-PT", "pt-PT": "pt-PT", "pt-pt": "pt-PT",
"fr": "fr", "fr-FR": "fr",
"es": "es", "es-ES": "es",
" zh ": "zh",
"": "",
"de": "de",
} {
if got := normalizePairLang(in); got != want {
t.Errorf("normalizePairLang(%q) = %q, want %q", in, got, want)
}
}
}
// French elision must yield its parts, or "j'ai" and "n'est" — two of the
// commonest shapes in the language — count for nothing.
func TestElisionYieldsMarkers(t *testing.T) {
if n := distinctMarkers("Je n'est pas", latinMarkers["fr"]); n < 3 {
t.Errorf("elided French: got %d markers, want >= 3 (je, est, pas)", n)
}
}
// Distinct, not total: one word repeated is one piece of evidence.
func TestRepeatedMarkerCountsOnce(t *testing.T) {
if n := distinctMarkers("que que que", latinMarkers["pt-PT"]); n != 1 {
t.Errorf("repeated marker: got %d, want 1", n)
}
}
+33
View File
@@ -201,3 +201,36 @@ func TestOfflineCardWinsSpanCollision(t *testing.T) {
t.Errorf("the exact offline card should own the span, got %+v", got[0])
}
}
// TestOfflineHanziFindingStaysMechanics: a 错别字 the Chinese rule pack found —
// both halves written in hanzi — files as an ordinary mechanics row.
//
// The check is worth its own test because there is a rule one layer over that
// would plausibly claim it. `isTranslation` re-labels an edit whose original
// reads as the writer's language and whose replacement reads as English, which
// is exactly how a zh-pair writer's quoted Chinese becomes a 'translate' card.
// A wrong-character fix looks like the first half of that and nothing like the
// second: 己经 → 已经 never leaves Chinese. It must stay a tidy-up in her own
// sentence, on the same rail as a doubled word, with no rendering-into-English
// implied anywhere.
func TestOfflineHanziFindingStaysMechanics(t *testing.T) {
srv, docID, _ := newTestServer(t, &stubClient{response: `{"suggestions":[]}`})
got := postMechanics(t, srv, docID, `[
{"from":1,"to":3,"original":"己经","replacement":"已经","explanation":"已经 (already) takes 已","type":"mechanics"}
]`)
if len(got) != 1 {
t.Fatalf("want the one finding, got %+v", got)
}
if got[0].Type != db.SuggestionTypeMechanics {
t.Errorf("hanzi fix filed as %q, want %q", got[0].Type, db.SuggestionTypeMechanics)
}
if got[0].Source != db.SuggestionSourceLocal {
t.Errorf("source = %q, want %q", got[0].Source, db.SuggestionSourceLocal)
}
// The characters survive the round trip intact — a mangled span here would
// replace the wrong characters in her document.
if got[0].Original != "己经" || got[0].Replacement != "已经" {
t.Errorf("round-tripped as %q → %q", got[0].Original, got[0].Replacement)
}
}
+303
View File
@@ -0,0 +1,303 @@
package suggestions
import (
"database/sql"
"gitea.parodia.dev/drwily/petal/internal/db"
"gitea.parodia.dev/drwily/petal/internal/llm"
)
// Reconciliation replaces the old "delete the family, insert the new batch"
// shape of every pass. A suggestion the pass proposes again is the *same*
// suggestion: it keeps its row, and therefore its id, its created_at and — most
// visibly — the explanation it was first given. The model re-words its reasoning
// every time it is asked, so re-inserting meant one unchanged mistake carried
// three different explanations in a single sitting.
//
// The id is what the frontend keys its cards on, so a stable id is also what
// keeps the rail from emptying and refilling, a card from collapsing mid-read,
// and the arrival chime from re-firing for advice she has already seen.
// pendingRow is the part of an existing pending suggestion reconciliation cares
// about.
type pendingRow struct {
id string
original string
replacement string
chunkHash string
from int
}
// loadPending reads the pending rows a pass owns. `where` is the pass's own
// scoping clause (by source, and for the model passes by family) — the same
// fragment that used to scope its DELETE.
func loadPending(tx *sql.Tx, docID, where string) ([]pendingRow, error) {
rows, err := tx.Query(
`SELECT id, original, replacement, chunk_hash, from_pos FROM suggestions
WHERE doc_id = ? AND status = ? AND `+where,
docID, db.SuggestionStatusPending,
)
if err != nil {
return nil, err
}
defer rows.Close()
var out []pendingRow
for rows.Next() {
var r pendingRow
if err := rows.Scan(&r.id, &r.original, &r.replacement, &r.chunkHash, &r.from); err != nil {
return nil, err
}
out = append(out, r)
}
return out, rows.Err()
}
// editKey identifies an edit by what it proposes, not where: "this exact change
// to this exact text". Normalized like the suppression comparisons, so the
// editor's quote rewriting and a reflowed paragraph don't read as a new edit.
func editKey(original, replacement string) string {
return normalizeForDedup(original) + "\x00" + normalizeForDedup(replacement)
}
// editIndex matches freshly proposed edits against the rows already standing.
type editIndex struct {
rows []pendingRow
used []bool
byKey map[string][]int
}
func indexByEdit(rows []pendingRow) *editIndex {
idx := &editIndex{rows: rows, used: make([]bool, len(rows)), byKey: map[string][]int{}}
for i, r := range rows {
k := editKey(r.original, r.replacement)
idx.byKey[k] = append(idx.byKey[k], i)
}
return idx
}
// take claims the standing row for this edit, if there is one. When a document
// repeats the same mistake, `near` (the fresh span's start) picks the closest
// standing row, so two identical cards keep their own identities instead of
// trading them whenever the text between them grows.
func (i *editIndex) take(original, replacement string, near int) (pendingRow, bool) {
best, bestDist := -1, 0
for _, n := range i.byKey[editKey(original, replacement)] {
if i.used[n] {
continue
}
d := i.rows[n].from - near
if d < 0 {
d = -d
}
if best < 0 || d < bestDist {
best, bestDist = n, d
}
}
if best < 0 {
return pendingRow{}, false
}
i.used[best] = true
return i.rows[best], true
}
// reposition updates the advisory offsets (and the sentence a row belongs to)
// without touching anything the writer can see. The frontend re-anchors by
// string at render time, so these only matter for the local-vs-model span
// arbitration in dedupeSpans.
func reposition(tx *sql.Tx, row pendingRow, from, to int, chunkHash string) error {
if row.from == from && row.chunkHash == chunkHash {
return nil
}
_, err := tx.Exec(
`UPDATE suggestions SET from_pos = ?, to_pos = ?, chunk_hash = ? WHERE id = ?`,
from, to, chunkHash, row.id,
)
return err
}
// reconcilePending brings a model pass's family in line with what it just
// proposed, sentence by sentence:
//
// - A row on a sentence this pass didn't ask about is kept untouched — that
// is the whole point of chunking. Only its offsets are refreshed.
// - A row on a sentence that no longer exists in the document is dropped: she
// rewrote or deleted it.
// - A row on a sentence the pass *did* ask about survives only if the model
// proposed the same edit again, in which case it keeps its identity.
//
// `fresh` names the sentences the model was asked about (nil when it wasn't
// called at all). inPlayAll marks the whole-document passes — voice and the
// collocation coach — where every row is up for re-proposal because the model
// just re-read everything.
//
// `pairLang` is the writer's own language, needed only to type a finding that
// turns out to be her language rendered into English (see language.go).
func (h *Handler) reconcilePending(
docID, contentText, pairLang string,
raw []llm.RawSuggestion,
scope pendingScope,
chunks, fresh []chunk,
inPlayAll bool,
) error {
tx, err := h.DB.Begin()
if err != nil {
return err
}
defer tx.Rollback()
existing, err := loadPending(tx, docID, scope.deleteWhere)
if err != nil {
return err
}
present := hashSet(chunks)
asked := hashSet(fresh)
modelRan := inPlayAll || fresh != nil
// Sentences to hand back to the model next time, because a row we were
// caching on them turned out to be unanchorable (see below).
reopen := map[string]bool{}
var inPlay []pendingRow
for _, r := range existing {
switch {
// A row whose sentence we can't name is never cached — it is re-examined
// whenever the model speaks, and left alone when it doesn't.
case inPlayAll, r.chunkHash == "" && modelRan, asked[r.chunkHash]:
inPlay = append(inPlay, r)
case r.chunkHash != "" && !present[r.chunkHash]:
if _, err := tx.Exec(`DELETE FROM suggestions WHERE id = ?`, r.id); err != nil {
return err
}
default:
// Untouched sentence: keep the card exactly as she last saw it.
from, to := locate(contentText, r.original)
if from < 0 {
// The sentence is unchanged in substance but the quoted span no
// longer matches byte for byte — a quote mark the editor rewrote
// inside it, say. The frontend anchors by that string, so this card
// can't be shown; drop it and let the sentence be read again rather
// than cache advice nobody can see.
reopen[r.chunkHash] = true
if _, err := tx.Exec(`DELETE FROM suggestions WHERE id = ?`, r.id); err != nil {
return err
}
continue
}
if err := reposition(tx, r, from, to, r.chunkHash); err != nil {
return err
}
}
}
for h := range reopen {
delete(present, h)
}
sup, err := buildSuppressor(tx, docID)
if err != nil {
return err
}
index := indexByEdit(inPlay)
kept := make(map[string]bool, len(inPlay))
for _, s := range raw {
if sup.suppressed(s.Original, s.Replacement) {
continue
}
from, to := locate(contentText, s.Original)
// Attribute the finding to a sentence the model was actually shown before
// falling back to the whole document: a short span ("the the") can occur in
// two sentences, and crediting it to the cached one would drop it as advice
// we already have.
hash := chunkFor(s.Original, fresh)
if hash == "" {
hash = chunkFor(s.Original, chunks)
}
// A sentence we didn't ask about already has whatever advice it deserves.
// The model can't normally quote one — it was only shown the delta — but if
// it wanders there anyway, the cached card stands rather than gaining a
// twin.
if !inPlayAll && hash != "" && present[hash] && !asked[hash] {
continue
}
if row, ok := index.take(s.Original, s.Replacement, from); ok {
kept[row.id] = true
if err := reposition(tx, row, from, to, hash); err != nil {
return err
}
continue
}
// A pass with a forced type owns its family outright and is never asked
// about translation: voice reads whole paragraphs for tone, and the
// collocation coach is about English word pairings. Only the open-typed
// grammar checkpoint can turn out to have been handed her own language.
typ := scope.forceType
if typ == "" {
typ = normalizeType(s.Type)
if isTranslation(s.Original, s.Replacement, pairLang) {
typ = db.SuggestionTypeTranslate
}
}
if _, err := tx.Exec(
`INSERT INTO suggestions (doc_id, from_pos, to_pos, original, replacement, explanation, type, source, chunk_hash)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)`,
docID, from, to, s.Original, s.Replacement, s.Explanation, typ, db.SuggestionSourceLLM, hash,
); err != nil {
return err
}
}
// Asked about and not proposed again: the model has changed its mind, or she
// has fixed it.
for _, r := range inPlay {
if kept[r.id] {
continue
}
if _, err := tx.Exec(`DELETE FROM suggestions WHERE id = ?`, r.id); err != nil {
return err
}
}
// Record the sentences this family has now read. Every sentence still in the
// document has been read by *some* pass: the ones just asked about now, the
// rest in an earlier round.
if scope.chunked {
if _, err := tx.Exec(
`DELETE FROM checked_chunks WHERE doc_id = ? AND family = ?`, docID, scope.family,
); err != nil {
return err
}
for h := range present {
if _, err := tx.Exec(
`INSERT INTO checked_chunks (doc_id, family, hash) VALUES (?, ?, ?)`,
docID, scope.family, h,
); err != nil {
return err
}
}
}
return tx.Commit()
}
// checkedChunks loads the sentences a family read on its last pass.
func (h *Handler) checkedChunks(docID, family string) (map[string]bool, error) {
rows, err := h.DB.Query(
`SELECT hash FROM checked_chunks WHERE doc_id = ? AND family = ?`, docID, family,
)
if err != nil {
return nil, err
}
defer rows.Close()
out := map[string]bool{}
for rows.Next() {
var hash string
if err := rows.Scan(&hash); err != nil {
return nil, err
}
out[hash] = true
}
return out, rows.Err()
}
+142
View File
@@ -0,0 +1,142 @@
package suggestions
import (
"encoding/json"
"net/http"
"testing"
"gitea.parodia.dev/drwily/petal/internal/db"
)
// The settled endpoint exists for the offline half of the loop. The rule pack
// detects from the text alone, 250 ms after a keystroke, and has no memory
// between runs — so without the document's record of what she has already
// answered, a dismissed finding is re-detected and re-rendered on the next
// keystroke, and stays there for as long as the server can't be reached.
func getSettled(t *testing.T, srv http.Handler, docID string) []string {
t.Helper()
rec := do(t, srv, http.MethodGet, "/docs/"+docID+"/settled", "")
if rec.Code != http.StatusOK {
t.Fatalf("settled: code=%d body=%s", rec.Code, rec.Body)
}
var out struct {
Originals []string `json:"originals"`
}
if err := json.Unmarshal(rec.Body.Bytes(), &out); err != nil {
t.Fatalf("decode: %v", err)
}
return out.Originals
}
func contains(list []string, want string) bool {
for _, s := range list {
if s == want {
return true
}
}
return false
}
// TestSettledListsActionedSpans proves the endpoint reports exactly the spans the
// suppressor would drop: accepted and dismissed, never pending. A pending row
// leaking in would be the damaging direction — the client would hide a card she
// has never been shown an answer to.
func TestSettledListsActionedSpans(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"I has","replacement":"I have","explanation":"agreement","type":"grammar"},
{"original":"two apple","replacement":"two apples","explanation":"plural","type":"grammar"}
]}`}
srv, docID, _ := newTestServer(t, client)
if got := getSettled(t, srv, docID); len(got) != 0 {
t.Fatalf("nothing actioned yet, got %v", got)
}
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var got []db.Suggestion
_ = json.Unmarshal(rec.Body.Bytes(), &got)
if len(got) != 2 {
t.Fatalf("first pass: want 2, got %d", len(got))
}
// One accepted, one still pending: only the accepted span is settled.
do(t, srv, http.MethodPost, "/suggestions/"+got[0].ID+"/accept", "")
settled := getSettled(t, srv, docID)
if len(settled) != 1 || settled[0] != got[0].Original {
t.Fatalf("want just %q settled, got %v", got[0].Original, settled)
}
// A dismissal settles a span just as an accept does — the whole point of the
// item: "you already decided about this one" doesn't mean "you agreed".
do(t, srv, http.MethodPost, "/suggestions/"+got[1].ID+"/dismiss", "")
settled = getSettled(t, srv, docID)
if len(settled) != 2 || !contains(settled, got[1].Original) {
t.Fatalf("dismissed span missing from %v", settled)
}
}
// TestSettledNormalizesAndDedupes proves the payload is normalized server-side
// and collapsed. The client compares its freshly-detected findings against these
// strings, so the two sides have to agree on what "the same span" is — the
// editor's quote churn is the case that breaks a byte-exact match, and it is why
// normalizeForDedup exists at all.
func TestSettledNormalizesAndDedupes(t *testing.T) {
client := &stubClient{response: `{"suggestions":[]}`}
srv, docID, h := newTestServer(t, client)
// The same span twice, differing only in quote style and line breaks — one
// accepted, one dismissed. Distinct rows; one settled span.
a := seedSuggestion(t, h, docID, "text", db.SuggestionTypeGrammar,
"She said \"hello\"\n to me", "She said 'hello' to me", "quotes")
b := seedSuggestion(t, h, docID, "text", db.SuggestionTypeGrammar,
"She said “hello” to me", "She said 'hello' to me", "quotes")
do(t, srv, http.MethodPost, "/suggestions/"+a+"/accept", "")
do(t, srv, http.MethodPost, "/suggestions/"+b+"/dismiss", "")
settled := getSettled(t, srv, docID)
if len(settled) != 1 {
t.Fatalf("two spellings of one span should collapse to one, got %v", settled)
}
if want := "She said 'hello' to me"; settled[0] != want {
t.Fatalf("settled[0] = %q, want normalized %q", settled[0], want)
}
}
// TestNormalizeMatchesTheClient is the Go half of a pair. Every case here also
// appears in web/src/lib/settled.test.ts, asserted against the TypeScript
// reimplementation of this function. The two are compared across a network
// boundary — the server normalizes what it sends, the client normalizes what it
// checks against it — so they have to fold the same characters the same way, and
// nothing but a shared list of cases can say so. Add to both or neither.
func TestNormalizeMatchesTheClient(t *testing.T) {
cases := []struct{ in, want string }{
{"She said “hello”", "She said 'hello'"},
{"She said \"hello\"", "She said 'hello'"},
{"its", "it's"},
{"its", "it's"},
{"`code´", "'code'"},
{" a apple\n here ", "a apple here"},
{"a\tapple", "a apple"},
{" \n ", ""},
{"我想说这句话", "我想说这句话"},
}
for _, c := range cases {
if got := normalizeForDedup(c.in); got != c.want {
t.Errorf("normalizeForDedup(%q) = %q, want %q", c.in, got, c.want)
}
}
}
// TestSettledEmptyIsAList guards the shape rather than the content: the client
// spreads this array into its settled set, and a null would throw there. Go
// marshals a nil slice as null, so this is one `[]string{}` away from breaking.
func TestSettledEmptyIsAList(t *testing.T) {
client := &stubClient{response: `{"suggestions":[]}`}
srv, docID, _ := newTestServer(t, client)
rec := do(t, srv, http.MethodGet, "/docs/"+docID+"/settled", "")
if body := rec.Body.String(); body != "{\"originals\":[]}\n" && body != "{\"originals\":[]}" {
t.Fatalf("empty settled body = %q, want an empty list", body)
}
}
+191
View File
@@ -0,0 +1,191 @@
package suggestions
import (
"encoding/json"
"net/http"
"strings"
"testing"
"gitea.parodia.dev/drwily/petal/internal/db"
"gitea.parodia.dev/drwily/petal/internal/llm"
)
// byOriginal indexes a pending set by the text each card flags.
func byOriginal(in []db.Suggestion) map[string]db.Suggestion {
out := map[string]db.Suggestion{}
for _, s := range in {
out[s.Original] = s
}
return out
}
// TestUntouchedSentencesKeepTheirCards is the heart of the stability work: she
// edits one sentence, and the cards on every other sentence stay exactly as they
// were — same id (so the rail keeps the card instead of remounting it), same
// explanation (the model re-words its reasoning every time it is asked, and one
// unchanged mistake used to carry three different explanations in a sitting).
// The model is only asked about the sentence that changed.
func TestUntouchedSentencesKeepTheirCards(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"I has two apple","replacement":"I have two apples","explanation":"first wording","type":"grammar"},
{"original":"She go to market","replacement":"She goes to market","explanation":"agreement","type":"grammar"}
]}`}
srv, docID, h := newTestServer(t, client)
h.Limit = llm.NewRateLimiter(0)
setDocText(t, h, docID, "I has two apple. She go to market yesterday.")
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var first []db.Suggestion
if err := json.Unmarshal(rec.Body.Bytes(), &first); err != nil {
t.Fatalf("decode: %v", err)
}
if len(first) != 2 {
t.Fatalf("first pass: want 2, got %d: %+v", len(first), first)
}
kept := byOriginal(first)["I has two apple"]
// She fixes only the second sentence. The model, asked again, re-words its
// reasoning about the first — which it must never get the chance to do.
setDocText(t, h, docID, "I has two apple. She goes to market yesterday.")
client.response = `{"suggestions":[
{"original":"I has two apple","replacement":"I have two apples","explanation":"REWORDED","type":"grammar"}
]}`
rec = do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var second []db.Suggestion
if err := json.Unmarshal(rec.Body.Bytes(), &second); err != nil {
t.Fatalf("decode: %v", err)
}
if strings.Contains(client.lastPrompt, "I has two apple") {
t.Fatalf("untouched sentence was sent to the model:\n%s", client.lastPrompt)
}
if !strings.Contains(client.lastPrompt, "She goes to market") {
t.Fatalf("edited sentence was not sent to the model:\n%s", client.lastPrompt)
}
now := byOriginal(second)["I has two apple"]
if now.ID != kept.ID {
t.Fatalf("card was remounted: id %q became %q", kept.ID, now.ID)
}
if now.Explanation != "first wording" {
t.Fatalf("explanation drifted: %q", now.Explanation)
}
// The fixed sentence's card is gone, and the model's stray re-proposal for the
// cached sentence did not become a second card.
if len(second) != 1 {
t.Fatalf("want exactly one card left, got %d: %+v", len(second), second)
}
}
// TestUnchangedDocumentSkipsTheModel proves a check with nothing new to read
// costs nothing: no model call, and every card left standing untouched. This is
// the doc-open and tone-less re-check path.
func TestUnchangedDocumentSkipsTheModel(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"I has","replacement":"I have","explanation":"agreement","type":"grammar"}
]}`}
srv, docID, h := newTestServer(t, client)
h.Limit = llm.NewRateLimiter(0)
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var first []db.Suggestion
_ = json.Unmarshal(rec.Body.Bytes(), &first)
if len(first) != 1 || client.calls != 1 {
t.Fatalf("first pass: %d cards, %d calls", len(first), client.calls)
}
rec = do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var second []db.Suggestion
if err := json.Unmarshal(rec.Body.Bytes(), &second); err != nil {
t.Fatalf("decode: %v", err)
}
if client.calls != 1 {
t.Fatalf("re-checking an unedited document called the model %d times", client.calls)
}
if len(second) != 1 || second[0].ID != first[0].ID {
t.Fatalf("card did not survive an idle re-check: %+v", second)
}
}
// TestDeletedSentenceDropsItsCard covers the other half of the skip path: she
// removes a flagged sentence outright, so nothing changed that the model could
// be asked about — but its card must still go.
func TestDeletedSentenceDropsItsCard(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"I has","replacement":"I have","explanation":"agreement","type":"grammar"}
]}`}
srv, docID, h := newTestServer(t, client)
h.Limit = llm.NewRateLimiter(0)
do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
setDocText(t, h, docID, "")
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
var got []db.Suggestion
if err := json.Unmarshal(rec.Body.Bytes(), &got); err != nil {
t.Fatalf("decode: %v", err)
}
if len(got) != 0 {
t.Fatalf("card outlived its sentence: %+v", got)
}
}
// TestToneChangeReopensEverySentence: the checkpoint's advice is written for the
// document's tone, so switching from a journal to an academic essay has to
// re-read sentences that haven't changed a character.
func TestToneChangeReopensEverySentence(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"I has","replacement":"I have","explanation":"agreement","type":"grammar"}
]}`}
srv, docID, h := newTestServer(t, client)
h.Limit = llm.NewRateLimiter(0)
do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
if _, err := h.DB.Exec(`UPDATE documents SET tone = 'academic' WHERE id = ?`, docID); err != nil {
t.Fatalf("set tone: %v", err)
}
do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
if client.calls != 2 {
t.Fatalf("tone change did not re-read the document: %d model calls", client.calls)
}
}
// TestMechanicsFindingsKeepTheirRows: the rule pack re-runs 250 ms after every
// keystroke. A finding it still reports must keep its row, or the rail would
// remount several times a sentence — collapsing a card she has open, and
// re-firing the arrival chime for advice she is already reading.
func TestMechanicsFindingsKeepTheirRows(t *testing.T) {
srv, docID, _ := newTestServer(t, &stubClient{response: `{"suggestions":[]}`})
body := `{"findings":[
{"from":0,"to":5,"original":"I has","replacement":"I have","explanation":"agreement"},
{"from":6,"to":15,"original":"two apple","replacement":"two apples","explanation":"plural"}
]}`
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/mechanics", body)
var first []db.Suggestion
if err := json.Unmarshal(rec.Body.Bytes(), &first); err != nil {
t.Fatalf("decode: %v", err)
}
if len(first) != 2 {
t.Fatalf("want 2 rows, got %d", len(first))
}
// She types elsewhere: same findings, shifted spans, one of them now fixed.
rec = do(t, srv, http.MethodPost, "/docs/"+docID+"/mechanics", `{"findings":[
{"from":20,"to":25,"original":"I has","replacement":"I have","explanation":"agreement"}
]}`)
var second []db.Suggestion
if err := json.Unmarshal(rec.Body.Bytes(), &second); err != nil {
t.Fatalf("decode: %v", err)
}
if len(second) != 1 {
t.Fatalf("want 1 row, got %d: %+v", len(second), second)
}
if second[0].ID != byOriginal(first)["I has"].ID {
t.Fatalf("surviving finding was given a new identity: %+v", second[0])
}
if second[0].FromPos != 20 {
t.Fatalf("span did not follow the text: %+v", second[0])
}
}
+107
View File
@@ -0,0 +1,107 @@
package suggestions
import (
"encoding/json"
"net/http"
"testing"
"gitea.parodia.dev/drwily/petal/internal/db"
)
// The pair model's flagship moment, end to end: she reaches for a sentence in
// her own language mid-document, and the card that comes back is labelled as a
// translation rather than as a tidy-up of her Chinese.
//
// The label is asserted through the real /check path rather than against
// isTranslation directly, because the point of the item was never the detector —
// Petal already found these spans and already rendered them into English. What
// was wrong was the type that reached the rail.
func TestChineseSpanBecomesATranslateCard(t *testing.T) {
// Note the model calls it "clarity", as the live build did. The type it
// volunteers is not consulted.
client := &stubClient{response: `{"suggestions":[
{"original":"我想说这句话但是不知道用英语怎么说。","replacement":"I want to say this but I don't know how to say it in English.","explanation":"这是英文说法 · Here is how to say it in English","type":"clarity"}
]}`}
srv, docID, database := newPairServer(t, client, "zh")
setDocTextDB(t, database, docID, "My weekend was good. 我想说这句话但是不知道用英语怎么说。")
var out []db.Suggestion
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
if rec.Code != http.StatusOK {
t.Fatalf("check: code=%d body=%s", rec.Code, rec.Body)
}
if err := json.Unmarshal(rec.Body.Bytes(), &out); err != nil {
t.Fatalf("decode: %v", err)
}
if len(out) != 1 {
t.Fatalf("want 1 card, got %d: %+v", len(out), out)
}
if out[0].Type != db.SuggestionTypeTranslate {
t.Fatalf("card type = %q, want %q", out[0].Type, db.SuggestionTypeTranslate)
}
// The rendering and the reasoning are the model's, untouched — only the label
// is Petal's.
if out[0].Replacement != "I want to say this but I don't know how to say it in English." {
t.Fatalf("replacement was rewritten: %q", out[0].Replacement)
}
}
// The other half of the same claim: an ordinary English correction on the same
// writer's document keeps the type the model gave it. A relabel that fired on
// everything would be no better than the label it replaced.
func TestEnglishCorrectionKeepsItsType(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"My weekend was very good","replacement":"My weekend was wonderful","explanation":"stronger wording","type":"phrasing"}
]}`}
srv, docID, database := newPairServer(t, client, "zh")
setDocTextDB(t, database, docID, "My weekend was very good.")
var out []db.Suggestion
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/check", "")
if err := json.Unmarshal(rec.Body.Bytes(), &out); err != nil {
t.Fatalf("decode: %v", err)
}
if len(out) != 1 {
t.Fatalf("want 1 card, got %d: %+v", len(out), out)
}
if out[0].Type != db.SuggestionTypePhrasing {
t.Fatalf("card type = %q, want %q", out[0].Type, db.SuggestionTypePhrasing)
}
}
// The voice pass reads whole paragraphs for tone and stamps its own family. A
// Chinese paragraph must not be able to smuggle a translate row into it — voice
// rows carry no replacement to accept, so a "translation" there would be a card
// offering nothing.
func TestVoicePassCannotProduceATranslateCard(t *testing.T) {
client := &stubClient{response: `{"suggestions":[
{"original":"我想说这句话但是不知道用英语怎么说。","replacement":"I want to say this in English.","explanation":"tone","type":"clarity"}
]}`}
srv, docID, database := newPairServer(t, client, "zh")
setDocTextDB(t, database, docID, "A first paragraph.\n\n我想说这句话但是不知道用英语怎么说。")
var out []db.Suggestion
rec := do(t, srv, http.MethodPost, "/docs/"+docID+"/voice", "")
if rec.Code != http.StatusOK {
t.Fatalf("voice: code=%d body=%s", rec.Code, rec.Body)
}
if err := json.Unmarshal(rec.Body.Bytes(), &out); err != nil {
t.Fatalf("decode: %v", err)
}
for _, s := range out {
if s.Type == db.SuggestionTypeTranslate {
t.Fatalf("voice pass produced a translate card: %+v", s)
}
}
}
// setDocTextDB is setDocText for the pair harness, which hands back the DB
// rather than the Handler.
func setDocTextDB(t *testing.T, database *db.DB, docID, text string) {
t.Helper()
if _, err := database.Exec(
`UPDATE documents SET content_text = ? WHERE id = ?`, text, docID,
); err != nil {
t.Fatalf("update doc text: %v", err)
}
}
+322
View File
@@ -0,0 +1,322 @@
#!/usr/bin/env python3
"""Build the two Chinese assets the learner direction of the zh pair needs.
Why two, and why they are split the way they are
------------------------------------------------
Every other pair Petal ships needs one asset: a word list the browser loads so
it can underline. Chinese needs two, because the browser and the server want
different halves of the same dictionary and for different reasons.
* **The browser needs a word list, and it needs it offline.** Chinese is
written without spaces, so there is no such thing as "the word under the
cursor" until something segments the sentence. Every ESL surface Petal
already has — the hover gloss, the right-click lookup, Ctrl/Cmd+D, the
vocabulary garden capture — is built on `wordAt`, and `wordAt` is a regex
over Latin letters. Segmentation is what replaces that regex, it runs on
every hover, and a round-trip per hover is not a hover. So the word list
ships to the browser: `web/public/dictionaries/zh/words.txt`.
* **The server holds the whole dictionary.** Pinyin and English senses are
only ever wanted one word at a time, in answer to a hover or a click, which
is exactly what `/api/gloss/{word}` already does for the other direction. So
the readings stay in the binary — `internal/lexicon/data/hanzi.json.gz` —
where their size costs a browser nothing.
That split is what makes the coverage decisions below come out *opposite* to
each other, and both are deliberate.
Two sources, because neither one has both halves
------------------------------------------------
* **CC-CEDICT** (CC BY-SA 4.0, https://www.mdbg.net/) has the headwords,
pinyin and English senses, and no frequency information at all.
* **jieba's `dict.txt`** (MIT, https://github.com/fxsjy/jieba) has ~349k
headwords with corpus frequencies, and no definitions.
Segmentation needs the frequencies: the standard algorithm is a shortest-path
walk over log-probabilities, not longest-match, and without frequencies the
classic ambiguities go the wrong way. The client list therefore carries
`word freq` per line; the gloss map carries readings.
The size decision is the client list, and it is a size decision only
--------------------------------------------------------------------
Measured on ordinary learner prose, the segmentation produced by the full jieba
dictionary (381,886 hanzi headwords once CC-CEDICT is unioned in) and by a
frequency-gated one is **identical**, including on the textbook ambiguities
(研究生命的起源, 乒乓球拍卖完了, 南京市长江大桥). What the long tail contains is
rare proper nouns, and the max-probability walk almost never chooses one: a
freq-3 name loses to two common words every time. The cases where a missing word
does change the answer degrade *gracefully* — the sentence splits into smaller
real words, which is a slightly clumsier gloss, not a wrong underline.
So the gate is set where the size is, at **freq >= 5**: 188,522 words, ~0.97 MB
gzipped over the wire, in line with fr (1.19 MB) and es (1.74 MB) rather than in
excess of them. Every CC-CEDICT headword is unioned back in regardless of
frequency, so the segmenter can always see a word the server can explain.
The gloss map is gated by nothing, for the opposite reason
-----------------------------------------------------------
The es phase settled that a *spelling* dictionary should hold the union of every
variety, because its only power is to underline and it must not underline
correct writing. This asset's only power is to **explain**, and the word a
learner stops on is precisely the one they do not know — which is to say, the
rare one. Trimming this by frequency would remove exactly the entries it exists
for. All 113,637 glossable headwords ship, ~3.1 MB gzipped, which is less than
half of what `synonyms.json.gz` has embedded since Phase 9.
Simplified only, and said out loud
-----------------------------------
The zh langpack is written in simplified characters and jieba's frequencies are
counted over simplified text, so the traditional headword in each CC-CEDICT line
is dropped and simplified is what both assets are keyed by. Glossing traditional
would be nearly free *here* and useless in the app: nothing would segment it, so
nothing would ever ask. Traditional support is a real feature and it starts with
a traditional word list, not with this file.
Usage:
curl -sL https://www.mdbg.net/chinese/export/cedict/cedict_1_0_ts_utf-8_mdbg.txt.gz | gunzip > cedict.txt
curl -sL https://raw.githubusercontent.com/fxsjy/jieba/master/jieba/dict.txt -o jieba.txt
python3 scripts/build_cedict.py cedict.txt jieba.txt \
web/public/dictionaries/zh/words.txt.gz \
internal/lexicon/data/hanzi.json.gz
"""
import gzip
import json
import re
import sys
# Frequency gate for the *client* list only (see the module docstring). Words
# below it survive if CC-CEDICT knows them, so "segmentable" is always a superset
# of "glossable" and a hover can never land on a word the server cannot explain.
MIN_FREQ = 5
# A CC-CEDICT headword we keep must be nothing but han characters. This drops the
# entries that are really English or numerals with a Chinese gloss attached
# ("AA制", "PM2.5", "11区"): the segmenter walks runs of hanzi, so a mixed
# headword can never be matched anyway, and a Latin one would collide with the
# English tokenizer that is still running on the same paragraph.
HANZI_ONLY = re.compile(r'^[一-鿿]+$')
CEDICT_LINE = re.compile(r'^(\S+) (\S+) \[(.*?)\] /(.*)/$')
# At most this many readings per word, and this many senses per reading. Two
# readings is not an arbitrary cap: it is what the particles need. 得 is dé "to
# obtain" *and* de, the complement marker — and a learner who hovers 得 in
# 说得很好 and is told only "to obtain" has been actively misinformed. Beyond two
# the tail is dialect and surnames, which crowd out the sense actually wanted.
MAX_READINGS = 2
MAX_SENSES = 3
MAX_SENSE_CHARS = 110
# Senses that describe the *dictionary* rather than the word. A learner hovering
# a word wants to know what it means, not that it is an orthographic variant of
# another headword they also do not know.
SKIP_SENSE_PREFIXES = ('variant of', 'old variant', 'see ', 'used in', 'abbr. for')
# ── pinyin: numbered syllables to tone marks ────────────────────────────────
# CC-CEDICT stores "gong1 yuan2". A learner reading their own writing back wants
# gōngyuán: the tone mark is the part that is hard to remember and the part that
# changes the word. The placement rule is the standard one — a/o/e take the mark
# if present, otherwise the last vowel of the final — and it is small enough to
# do here rather than to take a dependency for.
TONE_VOWELS = {
'a': 'āáǎà',
'e': 'ēéěè',
'i': 'īíǐì',
'o': 'ōóǒò',
'u': 'ūúǔù',
'ü': 'ǖǘǚǜ',
}
SYLLABLE = re.compile(r'^([a-zA-Zü:]+)([1-5])$')
def tone_mark(syllable: str) -> str:
"""One numbered pinyin syllable to its tone-marked form."""
m = SYLLABLE.match(syllable)
if not m:
# Punctuation, a bare letter (CC-CEDICT writes "X" for unknown), or an
# already-marked syllable: pass it through rather than mangling it.
return syllable
body, tone = m.group(1), int(m.group(2))
# CC-CEDICT writes ü as "u:" and, in a few entries, as "v".
body = body.replace('u:', 'ü').replace('U:', 'Ü').replace('v', 'ü').replace('V', 'Ü')
if tone == 5: # neutral tone carries no mark
return body
low = body.lower()
idx = -1
for vowel in ('a', 'o', 'e'):
idx = low.find(vowel)
if idx >= 0:
break
if idx < 0:
# No a/o/e: the mark goes on the last of i/u/ü (liú, guǐ, nǚ).
idx = max(low.rfind('i'), low.rfind('u'), low.rfind('ü'))
if idx < 0:
return body
marked = TONE_VOWELS[low[idx]][tone - 1]
if body[idx].isupper():
marked = marked.upper()
return body[:idx] + marked + body[idx + 1:]
def pinyin(numbered: str) -> str:
"""A whole CC-CEDICT pinyin field to tone marks, syllables joined up.
Joined rather than spaced because that is how a word is written when it is
being read as a word (gōngyuán, not gōng yuán); the spaces in the source are
a storage convention, not orthography.
"""
return ''.join(tone_mark(s) for s in numbered.split())
def clean_senses(raw: list[str]) -> list[str]:
"""Strip the apparatus CC-CEDICT carries for lexicographers, not learners."""
out = []
for sense in raw:
# "CL:座[zuo4]" is the measure-word field, useful and not a definition.
sense = re.sub(r'\s*CL:.*$', '', sense).strip()
# Bracketed pinyin cross-references ("abbr. for 的士[di1 shi4]").
sense = re.sub(r'\[[a-zA-Z0-9: ]+\]', '', sense).strip()
# Both edits cut inside parentheses — "cat (CL:只)" loses its closing
# bracket and leaves "cat (" on the card. Drop a dangling opener rather
# than trying to rebalance: what it introduced is gone.
if sense.count('(') > sense.count(')'):
sense = re.sub(r'\s*\([^()]*$', '', sense).strip()
if not sense or sense.startswith(SKIP_SENSE_PREFIXES):
continue
out.append(sense)
return out
def read_cedict(path: str) -> dict[str, list[tuple[str, list[str]]]]:
entries: dict[str, list[tuple[str, list[str]]]] = {}
for line in open(path, encoding='utf-8'):
if line.startswith('#'):
continue
m = CEDICT_LINE.match(line.strip())
if not m:
continue
_traditional, simplified, py, defs = m.groups()
if not HANZI_ONLY.match(simplified):
continue
entries.setdefault(simplified, []).append((py, defs.split('/')))
return entries
def read_jieba(path: str) -> dict[str, int]:
freqs: dict[str, int] = {}
for line in open(path, encoding='utf-8'):
parts = line.split()
if len(parts) >= 2 and HANZI_ONLY.match(parts[0]):
freqs[parts[0]] = int(parts[1])
return freqs
# ── the assertions ──────────────────────────────────────────────────────────
# The es phase's lesson, in the place it applies here: a check that every
# plausible input would pass is not a check. The Spanish MUST_ACCEPT list
# asserted vocabulary that all twenty-four builds carried, so it could not tell
# them apart. These assert the things that actually go wrong in *this* build —
# a mis-parsed pinyin field, a missing particle reading, a word list gated so
# hard the segmenter can no longer see a word the server can explain.
# Tone marking, including the three cases the placement rule exists for.
MUST_MARK = {
'gong1 yuan2': 'gōngyuán', # a/o/e rule, first syllable
'pao3 bu4': 'pǎobù',
'liu2': 'liú', # no a/o/e: mark the *last* of i/u
'gui3': 'guǐ',
'nu:3': '', # u: is ü
'lu:e4': 'lüè', # ü and an e in the same syllable: e wins
'de5': 'de', # neutral tone takes no mark at all
'Zhong1 wen2': 'Zhōngwén', # capitalised headword keeps its capital
}
# The particles the 错别字 rules are about must each carry the *grammatical*
# reading, not only the lexical one. 的/地/得 are the single most confused triple
# in written Chinese and all three are neutral-tone "de" in the use that matters;
# an entry that only knows 得 as dé is worse than no entry.
MUST_READ_DE = ('', '', '')
# Words the segmenter must be able to see. 图书馆 and 乒乓球 are ordinary
# vocabulary; 我 and 的 are the two commonest words in the language and a gate
# that dropped either would be visibly broken; 的士 is a CC-CEDICT headword rare
# enough to fall below the frequency gate, and is here to prove the union.
MUST_SEGMENT = ('', '', '图书馆', '乒乓球', '公园', '的士')
def check(words: dict[str, int], gloss: dict[str, list[list[str]]]) -> None:
for numbered, want in MUST_MARK.items():
got = pinyin(numbered)
assert got == want, f'pinyin({numbered!r}) = {got!r}, want {want!r}'
for particle in MUST_READ_DE:
readings = gloss.get(particle)
assert readings, f'{particle} has no gloss entry at all'
assert any(r[0] == 'de' for r in readings), \
f'{particle} never reads as neutral "de": {readings}'
for word in MUST_SEGMENT:
assert word in words, f'{word} missing from the segmentation list'
# The invariant the two gates exist to keep: everything the server can
# explain, the browser can find.
missing = [w for w in gloss if w not in words]
assert not missing, f'{len(missing)} glossable words are unsegmentable, e.g. {missing[:5]}'
# Nothing Latin leaked into either asset (see HANZI_ONLY).
for name, keys in (('words', words), ('gloss', gloss)):
bad = [k for k in keys if not HANZI_ONLY.match(k)]
assert not bad, f'non-hanzi headwords in {name}: {bad[:5]}'
def main() -> None:
if len(sys.argv) != 5:
sys.exit(__doc__.strip().rsplit('Usage:', 1)[-1].strip())
cedict_path, jieba_path, words_out, gloss_out = sys.argv[1:]
entries = read_cedict(cedict_path)
freqs = read_jieba(jieba_path)
# The client list: frequency-gated, then unioned with every glossable word.
# A CC-CEDICT word jieba has never seen gets frequency 1 — real, and rare
# enough that the max-probability walk will only choose it when nothing else
# fits, which is exactly the standing it should have.
words = {w: f for w, f in freqs.items() if f >= MIN_FREQ}
for w in entries:
words.setdefault(w, 1)
gloss: dict[str, list[list[str]]] = {}
for word, rows in entries.items():
readings: list[list[str]] = []
for numbered, defs in rows:
senses = clean_senses(defs)
if not senses:
continue
readings.append([pinyin(numbered), '; '.join(senses[:MAX_SENSES])[:MAX_SENSE_CHARS]])
if len(readings) == MAX_READINGS:
break
if readings:
gloss[word] = readings
check(words, gloss)
# Gzipped on disk, like the pt-PT/fr/es word lists: the browser inflates it
# with DecompressionStream (see useSpellChecker.fetchText), which costs no
# bundle bytes, and 0.97 MB over the wire rather than 2.23 MB is the whole
# difference between this and the biggest asset Petal ships.
body = ('\n'.join(f'{w} {words[w]}' for w in sorted(words)) + '\n').encode('utf-8')
with gzip.open(words_out, 'wb', compresslevel=9) as fh:
fh.write(body)
payload = json.dumps(gloss, ensure_ascii=False, separators=(',', ':')).encode('utf-8')
with gzip.open(gloss_out, 'wb', compresslevel=9) as fh:
fh.write(payload)
print(f'{words_out}: {len(words)} words, {len(body) / 1e6:.2f} MB raw, '
f'{len(gzip.compress(body, 9)) / 1e6:.2f} MB gzipped')
print(f'{gloss_out}: {len(gloss)} entries, {len(gzip.compress(payload, 9)) / 1e6:.2f} MB gzipped')
if __name__ == '__main__':
main()
+93
View File
@@ -89,6 +89,45 @@ through the packaging rather than through the model. The authentic dictionary is
the Projecto Natura one (Universidade do Minho) that LibreOffice ships and Debian
packages as `hunspell-pt-pt`; its aff declares `LANG pt_PT`.
**es: the wrong country again, hidden one layer further down.** Spanish looked
like it would repeat the pt trap — `hunspell-es` installs twenty country codes,
`es_AR` through `es_VE` — and then looked like it did not, because every one of
them is a symlink to a single `es_ES.aff`/`es_ES.dic`. Both readings were wrong.
Debian collapses the twenty because it ships **one** of upstream's builds, and
the one it ships is the **peninsular** `es_ES`. RLA (Santiago Bosio's project,
`sbosio/rla-es`) publishes twenty-four dictionaries per release: one per country,
plus a **generic `es`** that is the union of all of them. Debian packages neither
the generic one nor a choice — it packages Spain, under a name that reads like
"Spanish".
Measured against the v2.9 release: Debian's file is 659,085 expanded forms and
upstream `es_ES` is 659,018; the generic `es` is **717,640**. The 58,622-form
difference is almost entirely **voseo** — `vení`, `tenés`, `querés`, `sabés`,
`andá` — the present tense of most of Latin America, which Debian's package
rejects as misspellings. Petal ships the **generic** build.
**Vocabulary cannot detect this and morphology can.** The first version of the es
profile asserted the pan-Hispanic lexicon — *computadora* and *ordenador*, *papa*
and *patata* — and passed happily on the peninsular file, because **every** RLA
variant carries the full pan-Hispanic vocabulary; only the verb paradigms are
localised. The `REP` table is no help either: its `ll`↔`y` and `ás`↔`az` entries
look like evidence of yeísmo and seseo, but they are shared by all twenty-four
builds. What separates them is exactly two things, and the profile now demands
both at once: **voseo** (absent from `es_ES`) and **vosotros** (largely absent
from `es_MX`). Only the generic build has both, so only the generic build passes.
This is the same decision fr made between `-classical` and `-revised`, arriving
by a different road. The only thing this dictionary can do is underline
something, and *tienes* and *tenés* are both correct Spanish taught in different
countries — so Petal takes the build that accepts every variety rather than one
that makes a writer wrong for where she is from. Nothing is generated to get
there: the forms come from a real upstream package, which is what lets the
MUST_ACCEPT list prove which package it was.
Licensing note: RLA is tri-licensed GPL-3+ / LGPL-3+ / MPL-1.1+; Petal
redistributes under the MPL. The upstream README and LICENSE are vendored beside
the output.
**fr: the wrong side of an argument the French have not settled.** The regional
question turns out to be a non-question — Debian's `fr_FR`, `fr_CA`, `fr_BE`,
`fr_CH`, `fr_LU` and `fr_MC` are all symlinks to one `fr.dic`, so unlike pt there
@@ -114,6 +153,14 @@ Usage
src/usr/share/hunspell/fr.aff \\
src/usr/share/hunspell/fr.dic \\
web/public/dictionaries/fr
Spanish does not come from Debian — see below; `hunspell-es` is the peninsular
build. Take the generic dictionary from an upstream release instead:
curl -LO https://github.com/sbosio/rla-es/releases/download/v2.9/es.oxt
unzip -d src es.oxt # an .oxt is a zip
python3 scripts/build_hunspell_dictionary.py es \\
src/es.aff src/es.dic web/public/dictionaries/es
"""
import gzip
import os
@@ -404,6 +451,52 @@ PROFILES = {
"reject": ("jardinn", "écrivaitz", "xyzzyque"),
"wrong": "this does not look like the comprehensive French dictionary",
},
# The generic RLA build, and the accept list is written to reject the four
# neighbouring builds rather than to describe this one.
#
# The first version of this profile demanded *computadora* and *ordenador*,
# *papa* and *patata*, and passed — on the peninsular file, because **every**
# RLA variant carries the whole pan-Hispanic vocabulary. Vocabulary does not
# discriminate here at all; only morphology does, and it discriminates
# completely:
#
# * **voseo** (`vení`, `tenés`, `querés`) is in `es` and `es_AR` and not in
# `es_ES` or Debian's package. Demanding it rejects the peninsular build.
# * **vosotros** (`tenéis`, `escribid`) is in `es`, `es_AR` and `es_ES`, and
# largely absent from `es_MX`. Demanding it rejects the Mexican build.
#
# Requiring both at once leaves exactly one package standing: the generic
# `es`, which is the only one that accepts every variety of Spanish. That is
# the same reason fr ships `-comprehensive` — the only thing this dictionary
# can do is underline something, and *tienes* and *tenés* are both correct
# Spanish taught in different countries.
#
# The rest are shape checks: `escribiésemos` is the -se imperfect subjunctive,
# `dámelo` proves the enclitic pronoun rules ran, and `jardín`/`niño` prove
# FLAG UTF-8 was read as characters rather than bytes.
"es": {
"accept": (
# Rejects es_ES and Debian's hunspell-es.
"vení", "tenés", "querés", "sabés", "andá",
# Rejects es_MX.
"tenéis", "escribid",
# Rejects es_AR, which has both voseo and vosotros and would
# otherwise pass. Caribbean and Andean everyday words: the generic
# build is the union of all twenty-four, so it is the only one that
# holds another region's vocabulary as well as its own.
"arepa", "chévere", "bacán",
# Pan-Hispanic vocabulary. These pass on every RLA build, so they
# prove nothing on their own — kept because a source that stopped
# being RLA at all would fail them.
"computadora", "ordenador", "papa", "patata", "jugo", "zumo",
# Morphology and encoding.
"escribiéramos", "escribiésemos", "escríbeme", "dámelo",
"jardín", "niño", "corazón",
),
"reject": ("jardinn", "escribiz", "xyzzyque", "haiga"),
"wrong": "this is not the generic RLA build (a per-country one accepts "
"only some of these)",
},
}
+18
View File
@@ -30,6 +30,7 @@
},
"devDependencies": {
"@tailwindcss/vite": "^4.0.0",
"@types/node": "^26.1.2",
"@types/react": "^19.1.0",
"@types/react-dom": "^19.1.0",
"@vitejs/plugin-react": "^4.3.4",
@@ -2107,6 +2108,16 @@
"integrity": "sha512-RGdgjQUZba5p6QEFAVx2OGb8rQDL/cPRG7GiedRzMcJ1tYnUANBncjbSB1NRGwbvjcPeikRABz2nshyPk1bhWg==",
"license": "MIT"
},
"node_modules/@types/node": {
"version": "26.1.2",
"resolved": "https://registry.npmjs.org/@types/node/-/node-26.1.2.tgz",
"integrity": "sha512-Vu4a5UFA9rIIFJ7rB/Vaafh9lrCQszopTCx6KjFboXTGQbPNasehVR5TEiithSDGyd1DEiUByggTZsg8jukeIg==",
"dev": true,
"license": "MIT",
"dependencies": {
"undici-types": "~8.3.0"
}
},
"node_modules/@types/react": {
"version": "19.2.17",
"resolved": "https://registry.npmjs.org/@types/react/-/react-19.2.17.tgz",
@@ -3559,6 +3570,13 @@
"integrity": "sha512-ARDJmphmdvUk6Glw7y9DQ2bFkKBHwQHLi2lsaH6PPmz/Ka9sFOBsBluozhDltWmnv9u/cF6Rt87znRTPV+yp/A==",
"license": "MIT"
},
"node_modules/undici-types": {
"version": "8.3.0",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-8.3.0.tgz",
"integrity": "sha512-j375ScV60dom+YkPFIfTLcOiPxkN/buHz5GobjLhixFuANaNs3C9l4GmrWqejgXWJ7BbJcFYpTEUkS1Ge8bpZQ==",
"dev": true,
"license": "MIT"
},
"node_modules/update-browserslist-db": {
"version": "1.2.3",
"resolved": "https://registry.npmjs.org/update-browserslist-db/-/update-browserslist-db-1.2.3.tgz",
+1
View File
@@ -33,6 +33,7 @@
},
"devDependencies": {
"@tailwindcss/vite": "^4.0.0",
"@types/node": "^26.1.2",
"@types/react": "^19.1.0",
"@types/react-dom": "^19.1.0",
"@vitejs/plugin-react": "^4.3.4",
+68
View File
@@ -0,0 +1,68 @@
Spanish spelling dictionary
===========================
The word list in `es.dic.gz` and the suggestion directives in `es.aff` are
derived from the **generic** Spanish Hunspell dictionary published by the RLA-ES
project ("Recursos Lingüísticos Abiertos del Español"), release v2.9.
Copyright (C) Santiago Bosio and the RLA-ES contributors
License: GPL-3+ or LGPL-3+ or MPL-1.1+
Tri-licensed; you may choose freely among the three. Petal
redistributes under the MPL. Full texts:
https://www.gnu.org/licenses/gpl-3.0.en.html
https://www.gnu.org/licenses/lgpl-3.0.en.html
https://www.mozilla.org/en-US/MPL/1.1/
Upstream: https://github.com/sbosio/rla-es
Source: https://github.com/sbosio/rla-es/releases/download/v2.9/es.oxt
(an .oxt is a zip; es.aff and es.dic are at its root)
Not the Debian package, and that is the point
---------------------------------------------
`hunspell-es` looks like the obvious source and is the wrong one. It installs
twenty country codes — `es_AR` through `es_VE` — all symlinked to a single file,
which reads like "one pan-Hispanic dictionary". It is not. RLA publishes
twenty-four dictionaries per release: one per country, plus a **generic `es`**
that is the union of all of them, and Debian ships the **peninsular `es_ES`**
build under the collapsed name.
Measured against v2.9, expanded to surface forms:
Debian hunspell-es 659,085 forms voseo: no vosotros: yes
upstream es_ES 659,018 forms voseo: no vosotros: yes
upstream es_MX 554,923 forms voseo: no vosotros: no
upstream es_AR 669,605 forms voseo: yes vosotros: yes
upstream es (generic) 717,640 forms voseo: yes vosotros: yes <-- this
The 58,622-form gap between Debian's file and the generic one is essentially the
**voseo** paradigm — `vení`, `tenés`, `querés`, `sabés`, `andá` — the ordinary
present tense of Argentina, Uruguay, Paraguay and much of Central America. Under
the Debian package, a writer using it would have had her own verbs underlined as
misspellings.
Why the generic build rather than one country
---------------------------------------------
The only thing this dictionary can do is underline something. *Tienes* and
*tenés* are both correct Spanish, taught in different countries, and a writing
companion has no business marking one of them wrong — the same reasoning that
makes the French dictionary here the `-comprehensive` packaging rather than
`-classical` or `-revised`. The generic build accepts every variety, so Petal
underlines only what no Spanish speaker anywhere would write.
How the build proves it got this file
-------------------------------------
Vocabulary cannot tell these builds apart: *every* RLA variant carries the full
pan-Hispanic lexicon, so *computadora* alongside *ordenador* passes on the
peninsular file too. (The `REP` table is likewise no evidence — its `ll`/`y` and
`ás`/`az` entries look like yeísmo and seseo but are shared by all builds.) Only
the verb paradigms are localised, so the `es` profile in
`scripts/build_hunspell_dictionary.py` demands, all at once:
* **voseo** (`vení`, `tenés`, `querés`) — rejects `es_ES` and Debian's package;
* **vosotros** (`tenéis`, `escribid`) — rejects `es_MX`;
* **another region's everyday words** (`arepa`, `chévere`, `bacán`) — rejects
`es_AR`, which has both paradigms and would otherwise pass.
Only the generic build satisfies all three. Each of the four neighbouring builds
was run through the profile and confirmed to fail.
+29
View File
@@ -0,0 +1,29 @@
SET UTF-8
TRY aeroinsctldumpbgfvhzóíjáqéñxyúükwAEROINSCTLDUMPBGFVHZÓÍJÁQÉÑXYÚÜKW
REP 19
REP ás az
REP az ás
REP cc x
REP és ez
REP ez és
REP güe hue
REP güi hui
REP hue güe
REP hui güi
REP ís iz
REP ío ido
REP ke que
REP ki qui
REP ll y
REP mb nv
REP nv mb
REP seci cesi
REP x cc
REP y ll
MAP 6
MAP aáAÁ
MAP eéEÉ
MAP iíIÍ
MAP oóOÓ
MAP uúüUÚÜ
MAP nñNÑ
Binary file not shown.
+59
View File
@@ -0,0 +1,59 @@
Chinese word list (segmentation)
================================
`words.txt.gz` is not a spelling dictionary — Chinese has no spelling to check
in the Hunspell sense. It is the word list Petal's segmenter walks, so that a
sentence written without spaces has words in it to hover, look up and capture.
Each line is `word frequency`. See scripts/build_cedict.py for how it is built
and why it is gated where it is.
It is derived from two upstream sources, both redistributable, both credited
here because the file itself has no room for a header.
CC-CEDICT — the headwords
-------------------------
Community maintained free Chinese-English dictionary, published by MDBG.
https://www.mdbg.net/chinese/dictionary?page=cedict
Licensed under the Creative Commons Attribution-ShareAlike 4.0 International
License — https://creativecommons.org/licenses/by-sa/4.0/
Referenced works:
CEDICT — Copyright (C) 1997, 1998 Paul Andrew Denisowski
CC-CEDICT is also the source of `internal/lexicon/data/hanzi.json.gz`, the
pinyin and English senses embedded in the Petal binary. The same attribution and
the same ShareAlike terms apply to that file; it is named here because it has
nowhere of its own to say so.
jieba — the frequencies
-----------------------
"结巴" Chinese word segmentation, by Sun Junyi.
https://github.com/fxsjy/jieba
MIT License
Copyright (c) 2013 Sun Junyi
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
Only the word/frequency columns are used; jieba's part-of-speech tags and its
algorithm are not (Petal's segmenter is its own, in web/src/lib/segment.ts).
Binary file not shown.
+81 -16
View File
@@ -1,8 +1,9 @@
import { useCallback, useEffect, useRef, useState } from 'react'
import { api, type DocSummary, type DocUpdate, type Document, type Suggestion, type Tag, type TagColor } from './api/client'
import { useAutoSave } from './hooks/useAutoSave'
import { useCheckpoint } from './hooks/useCheckpoint'
import { findingKey, useCheckpoint } from './hooks/useCheckpoint'
import { useSpellChecker } from './hooks/useSpellChecker'
import { useSegmenter } from './hooks/useSegmenter'
import { useTags } from './hooks/useTags'
import { DocList } from './components/DocList/DocList'
import { EditorCore, type EditorChange } from './components/Editor/EditorCore'
@@ -22,6 +23,7 @@ import { PetalFall } from './effects/PetalFall'
import { usePack } from './i18n'
import { useNightMode } from './hooks/useNightMode'
import { playSuggestionSound } from './audio/sounds'
import { fromIME } from './lib/ime'
export default function App() {
const updateAvailable = useVersionWatch()
@@ -31,7 +33,7 @@ export default function App() {
const night = useNightMode()
// Who's writing, and whether the server still recognises them. `signedOut`
// flips the moment any call comes back 401.
const { me, signedOut } = useSession()
const { me, signedOut, setDirection, setPair } = useSession()
const t = usePack()
// A real account to sign out of, as opposed to the hardcoded local user a
// build without auth configured runs as.
@@ -80,6 +82,13 @@ export default function App() {
}, [])
const { status, schedule, saveNow } = useAutoSave(currentDoc?.id ?? null)
// The Chinese word list, for a writer going the other way through the zh pair.
// Gated on the account's own setting rather than on anything in the text: a
// Mandarin native drafting English quotes Chinese constantly, and none of that
// is what segmentation is for. Declared above the checkpoint because the
// offline 错别字 pass reads it.
const segmenter = useSegmenter(me?.direction === 'learning_pair')
const {
suggestions,
checking,
@@ -90,7 +99,8 @@ export default function App() {
runVoice,
runCollocation,
removeSuggestion,
} = useCheckpoint(currentDoc?.id ?? null)
resolveServerId,
} = useCheckpoint(currentDoc?.id ?? null, segmenter)
// Browser-side spell checker — loads the en-US dictionary once per session.
const { checker: spellChecker, addWord } = useSpellChecker()
// The tag roster (with counts). Assignments live on the doc summaries below.
@@ -322,14 +332,25 @@ export default function App() {
const handleEditorChange = useCallback(
(change: EditorChange) => {
const { composing, ...patch } = change
setWordCount(change.word_count)
setDocText(change.content_text)
setEditTick((n) => n + 1)
if (currentDoc) {
patchSummary(currentDoc.id, { word_count: change.word_count })
schedule(change)
scheduleCheckpoint(change.content_text)
// The save is never held: see EditorChange.composing. The flag itself
// stays out of the patch — it describes the keyboard, not the document,
// and the stashed draft a signed-out save leaves behind should be the
// document alone.
schedule(patch)
}
// Everything below reads the text as prose. While an IME composition is
// in flight it is not prose yet — it is the pinyin she is converting — so
// the checkpoint, the rule pack and the companion all wait for the word
// to commit. EditorCore emits one more change the moment it does, so
// nothing is skipped, only deferred by the length of a word.
if (composing) return
setDocText(change.content_text)
setEditTick((n) => n + 1)
if (currentDoc) scheduleCheckpoint(change.content_text)
},
[currentDoc, patchSummary, schedule, scheduleCheckpoint],
)
@@ -350,17 +371,44 @@ export default function App() {
// Accept applies the replacement in the editor (handled in EditorCore) and
// marks the suggestion accepted; dismiss just rejects it. Both drop it locally.
// A rule-pack card can be accepted before its row exists — the edit has already
// landed either way, so a missing id just means there's nothing to file.
const handleAccept = useCallback(
async (s: Suggestion) => {
removeSuggestion(s.id)
setAcceptTick((n) => n + 1)
try {
await api.acceptSuggestion(s.id)
const id = await resolveServerId(s)
if (id) await api.acceptSuggestion(id)
} catch (err) {
console.error('accept failed', err)
}
},
[removeSuggestion],
[removeSuggestion, resolveServerId],
)
// Accept-all: EditorCore has already applied the whole category in one editor
// transaction, so this is only the bookkeeping. Each row is filed individually
// (there's no batch endpoint, and each accept plants its own word in the
// garden), but the kitten cheers once — five cheers for one click would read as
// five separate congratulations for a decision she made once.
const handleAcceptMany = useCallback(
async (list: Suggestion[]) => {
if (list.length === 0) return
for (const s of list) removeSuggestion(s.id)
setAcceptTick((n) => n + 1)
await Promise.all(
list.map(async (s) => {
try {
const id = await resolveServerId(s)
if (id) await api.acceptSuggestion(id)
} catch (err) {
console.error('accept failed', err)
}
}),
)
},
[removeSuggestion, resolveServerId],
)
// After restoring a version, swap the restored doc into the editor. Bumping
@@ -378,11 +426,14 @@ export default function App() {
[patchSummary],
)
// Escape always restores the sidebar while in distraction-free mode.
// Escape always restores the sidebar while in distraction-free mode — unless
// it belongs to an IME, where it cancels a candidate and never reaches Petal
// at all. This is the writer typing Chinese in the very mode built for
// uninterrupted writing, so it is the one worth getting right.
useEffect(() => {
if (!focusMode) return
const onKey = (e: KeyboardEvent) => {
if (e.key === 'Escape') setFocusMode(false)
if (e.key === 'Escape' && !fromIME(e)) setFocusMode(false)
}
window.addEventListener('keydown', onKey)
return () => window.removeEventListener('keydown', onKey)
@@ -398,12 +449,13 @@ export default function App() {
async (s: Suggestion) => {
removeSuggestion(s.id)
try {
await api.dismissSuggestion(s.id)
const id = await resolveServerId(s)
if (id) await api.dismissSuggestion(id)
} catch (err) {
console.error('dismiss failed', err)
}
},
[removeSuggestion],
[removeSuggestion, resolveServerId],
)
// Play a soft sound when freshly-checked suggestions arrive — one per distinct
@@ -411,10 +463,14 @@ export default function App() {
// a pile-up. We track which ids we've already chimed for, and only chime for
// recently-created suggestions so opening a doc with old pending advice stays
// silent (the existing set was created in a past session).
// Rule-pack findings are chimed by their wording, not their id: the same fix
// appears first as a provisional card and then as its persisted row, and the
// writer should hear it once.
const chimedRef = useRef<Set<string>>(new Set())
useEffect(() => {
const fresh = suggestions.filter((s) => !chimedRef.current.has(s.id))
fresh.forEach((s) => chimedRef.current.add(s.id))
const key = (s: Suggestion) => (s.source === 'local' ? `local:${findingKey(s)}` : s.id)
const fresh = suggestions.filter((s) => !chimedRef.current.has(key(s)))
fresh.forEach((s) => chimedRef.current.add(key(s)))
const justMade = fresh.filter(
(s) => Date.now() - new Date(s.created_at).getTime() < 12_000,
)
@@ -495,6 +551,9 @@ export default function App() {
onToggleTag={handleToggleTag}
onCreateTag={handleCreateTag}
account={account}
direction={me?.direction}
onDirection={setDirection}
onPair={setPair}
/>
</div>
@@ -509,7 +568,10 @@ export default function App() {
<>
<div
onMouseDown={handleChromeDown}
className="flex flex-1 flex-col overflow-y-auto px-6 py-8"
// `petal-scrollport` marks this as the editor's scrolling
// ancestor; EditorCore measures it to pin the text column while
// the suggestion rail's overhang is scrolled.
className="petal-scrollport flex flex-1 flex-col overflow-y-auto px-6 py-8"
>
<div ref={canvasRef} className="mx-auto flex w-full max-w-[720px] flex-1 flex-col">
{/* Title, then the three chrome pills. Their labels are
@@ -555,8 +617,10 @@ export default function App() {
docId={currentDoc.id}
initialContent={currentDoc.content}
onChange={handleEditorChange}
segmenter={segmenter}
suggestions={suggestions}
onAccept={handleAccept}
onAcceptMany={handleAcceptMany}
onDismiss={handleDismiss}
onVoiceCheck={runVoice}
voicing={voicing}
@@ -577,6 +641,7 @@ export default function App() {
voicing={voicing}
collocating={collocating}
llmDown={llmDown}
suggestionCount={suggestions.length}
/>
</div>
</>
+53 -1
View File
@@ -101,7 +101,18 @@ export interface Gloss {
reverse?: string
}
export type SuggestionType = 'grammar' | 'phrasing' | 'idiom' | 'clarity' | 'voice' | 'collocation' | 'mechanics'
// 'translate' is a span she wrote in her own language, rendered into English —
// not a correction. The server decides the label from the span itself, never from
// the model, so the client can trust it (see suggestions/language.go).
export type SuggestionType =
| 'grammar'
| 'phrasing'
| 'idiom'
| 'clarity'
| 'translate'
| 'voice'
| 'collocation'
| 'mechanics'
// One word in the vocabulary garden: a looked-up word with its gloss/phonetic,
// the sentence it was met in, and its spaced-repetition state. `reps` drives how
@@ -199,6 +210,21 @@ export interface PersonalWords {
words: string[]
}
// One pronunciation of a Chinese word, and what it means in that pronunciation.
// A list, because 得 is dé "to obtain" and also the particle in 说得很好.
export interface HanziReading {
pinyin: string
senses: string
}
// A Chinese word lookup. `readings` is empty for a word with no headword, in
// which case `chars` may carry the character-by-character reading.
export interface HanziInfo {
word: string
readings: HanziReading[]
chars: { char: string; pinyin: string; senses: string }[]
}
// Who's writing. Mirrors the backend db.User.
export interface Me {
id: string
@@ -206,6 +232,11 @@ export interface Me {
display_name: string
created_at: string
pair_lang: string
// Which half of the pair is being learned: 'learning_en' (the writer is
// native in pair_lang and practising English) or 'learning_pair' (the other
// way round). Mirrors users.direction; the server refuses 'learning_pair' for
// a pair it has no word list for.
direction: string
}
// Thrown when the server says the session is gone. Callers can tell it apart
@@ -258,6 +289,15 @@ export const api = {
setPairLang: (lang: string) =>
req<Me>('/me', { method: 'PATCH', body: JSON.stringify({ pair_lang: lang }) }),
// Turn the pair around. Same endpoint, same contract, and deliberately a
// separate call: the two fields are validated together server-side, so a
// client that wants to change both says both in one request rather than
// sending two that each pass on their own.
setDirection: (direction: string) =>
req<Me>('/me', { method: 'PATCH', body: JSON.stringify({ direction }) }),
setPair: (lang: string, direction: string) =>
req<Me>('/me', { method: 'PATCH', body: JSON.stringify({ pair_lang: lang, direction }) }),
listDocs: () => req<DocSummary[]>('/docs'),
createDoc: () => req<Document>('/docs', { method: 'POST' }),
getDoc: (id: string) => req<Document>(`/docs/${id}`),
@@ -287,6 +327,12 @@ export const api = {
}),
// Pending suggestions for a doc, loaded when the editor opens it.
listSuggestions: (id: string) => req<Suggestion[]>(`/docs/${id}/suggestions`),
// The spans she has already accepted or dismissed on this doc, normalized. The
// server suppresses these itself; the client needs them so the instant rule-pack
// pass doesn't hand back a dismissed card before the server can say otherwise —
// or, with the server unreachable, at all. See lib/settled.ts.
listSettled: (id: string) =>
req<{ originals: string[] }>(`/docs/${id}/settled`),
acceptSuggestion: (id: string) =>
req<void>(`/suggestions/${id}/accept`, { method: 'POST' }),
dismissSuggestion: (id: string) =>
@@ -331,6 +377,12 @@ export const api = {
// Lightweight Chinese-only gloss for the inline hover/select tooltip — instant
// and offline, so it fires on hover without spinning up the heavier lookup.
glossWord: (word: string) => req<Gloss>(`/gloss/${encodeURIComponent(word)}`),
// The same lookup pointing the other way: a Chinese word to its pinyin and
// English senses, for an account learning the pair language rather than
// English. A word the dictionary has no headword for comes back with empty
// readings and — when its characters are known — a per-character reading
// instead, which is a real second answer for a compound.
hanziWord: (word: string) => req<HanziInfo>(`/hanzi/${encodeURIComponent(word)}`),
// Tone-rewrite: rewrites a selected passage in the given style ('natural',
// 'academic', …) and returns the rewritten text for an in-editor preview. Not
// persisted — the editor applies it directly on accept.
@@ -3,6 +3,7 @@ import type { SaveStatus } from '../../hooks/useAutoSave'
import { useCompanion, type Mood } from './useCompanion'
import { LottiePlayer } from './LottiePlayer'
import { COMPANIONS, DEFAULT_COMPANION } from './companions'
import { useCardOverlap } from './useCardOverlap'
import { onPrefsScopeChange, readPref, writePref } from '../../lib/prefs'
import { usePack } from '../../i18n'
@@ -86,6 +87,15 @@ export function PetalCompanion({
const companion = COMPANIONS.find((c) => c.id === companionId) ?? COMPANIONS[0]
const [pickerOpen, setPickerOpen] = useState(false)
const rootRef = useRef<HTMLDivElement>(null)
const badgeRef = useRef<HTMLButtonElement>(null)
// When suggestion cards stack down into the corner, the kitten fades to
// translucent and shrinks a step so the card stays readable and clickable.
// It wakes back up whenever it has something to say (bubble) or is being
// interacted with (picker open) — except under an open History or Garden
// panel, where even a cheer would cover the controls she just reached for.
const crowded = useCardOverlap(badgeRef)
const faded = crowded.modal || (crowded.cards && !pickerOpen && !bubble)
// Awake companions (no sleeping clip) don't visibly nap — when the engine
// dozes them, keep their normal idle pose instead of a sleepy face. Only a
@@ -176,7 +186,10 @@ export function PetalCompanion({
</div>
)}
{bubble && !pickerOpen && (
{/* The bubble is its own layer, so fading the badge doesn't hide it —
hold it back explicitly while a panel is open. useCompanion keeps the
bubble in state, so it reappears when she closes the panel. */}
{bubble && !pickerOpen && !crowded.modal && (
<div
role="status"
onClick={dismiss}
@@ -239,11 +252,14 @@ export function PetalCompanion({
)}
<button
ref={badgeRef}
type="button"
onClick={() => setPickerOpen((o) => !o)}
title="Choose a companion"
aria-label="Choose a companion"
className={`petal-companion pointer-events-auto select-none ${napping ? 'petal-companion-sleep' : ''}`}
className={`petal-companion select-none ${faded ? 'petal-companion-faded' : 'pointer-events-auto'}${
napping ? ' petal-companion-sleep' : ''
}`}
style={{
// Size scales with the viewport — see --petal-companion-size in index.css.
width: 'var(--petal-companion-size)',
+141
View File
@@ -0,0 +1,141 @@
import { readFileSync } from 'node:fs'
import { gunzipSync } from 'node:zlib'
import { describe, expect, it } from 'vitest'
import { CONFUSION_PAIRS, hanziFindings } from './hanzi'
import { buildSegmenter } from '../../lib/segment'
// The 错别字 pack, held to the bar Phase 22 set for the English rule pack: every
// rule pinned in *two* directions — the mistake it must catch, and the correct
// writing next to it that it must leave alone.
//
// Here the second direction is the one that matters, and it is unusually easy to
// get wrong. Chinese has no spaces, so every one of these rules is a substring
// match on running text, and for most of them there exists an ordinary correct
// sentence that contains the substring across a word boundary. Those sentences
// are the real test.
const raw = gunzipSync(readFileSync(new URL('../../../public/dictionaries/zh/words.txt.gz', import.meta.url)))
const seg = buildSegmenter(raw.toString('utf8'))
const flagged = (text: string) => hanziFindings(text, seg).map((f) => `${f.original}${f.replacement}`)
describe('the gate that admits a rule', () => {
// The pack's own claim about itself, checked against the shipped dictionary
// rather than asserted in a comment. A pair whose wrong form is a real word
// cannot be decided mechanically and does not belong here.
it('every wrong form is not a word, and every right form is', () => {
for (const { wrong, right } of CONFUSION_PAIRS) {
expect(seg.has(wrong), `${wrong} is a dictionary word and must not be flagged`).toBe(false)
expect(seg.has(right), `${right} is not a dictionary word`).toBe(true)
}
})
// The errors this pack deliberately refuses, and why — each is a genuine
// mistake by a modern standard whose wrong form is itself a headword. If a
// dictionary rebuild ever drops one of these, this test fails and the pair
// becomes admissible; that is the intended way to find out.
it('refuses the well-known errors it cannot decide', () => {
for (const undecidable of ['自已', '好象', '倒底', '帐号', '部份']) {
expect(seg.has(undecidable), `${undecidable} is no longer a word — reconsider the rule`).toBe(true)
expect(flagged(`这是${undecidable}的例子`)).toEqual([])
}
})
})
describe('the mistakes it catches', () => {
it('已 / 己 / 以', () => {
expect(flagged('我己经写完了作业')).toEqual(['己经→已经'])
expect(flagged('我以经吃过饭了')).toEqual(['以经→已经'])
expect(flagged('下课已后我们去公园')).toEqual(['已后→以后'])
})
it('在 / 再', () => {
expect(flagged('明天在见')).toEqual(['在见→再见'])
expect(flagged('他正再看书')).toEqual(['正再→正在'])
expect(flagged('现再几点了')).toEqual(['现再→现在'])
})
it('做 / 作', () => {
expect(flagged('我的工做很忙')).toEqual(['工做→工作'])
expect(flagged('老师给我们很多做业')).toEqual(['做业→作业'])
expect(flagged('这本书的做者是谁')).toEqual(['做者→作者'])
})
it('the rest', () => {
expect(flagged('我觉的这个很好')).toEqual(['觉的→觉得'])
expect(flagged('你因该早点睡')).toEqual(['因该→应该'])
expect(flagged('即然你来了就坐下吧')).toEqual(['即然→既然'])
expect(flagged('你知到吗')).toEqual(['知到→知道'])
expect(flagged('请输入你的蜜码')).toEqual(['蜜码→密码'])
})
it('reports an exact span, so the card replaces the right characters', () => {
const text = '我己经到了'
const [f] = hanziFindings(text, seg)
expect(text.slice(f.from, f.to)).toBe('己经')
expect(text.slice(0, f.from) + f.replacement + text.slice(f.to)).toBe('我已经到了')
})
it('finds every occurrence, in document order', () => {
expect(flagged('我己经吃了,他也己经吃了')).toEqual(['己经→已经', '己经→已经'])
expect(flagged('我的工做很忙,所以我觉的很累')).toEqual(['工做→工作', '觉的→觉得'])
})
})
// ── the direction that matters ──────────────────────────────────────────────
describe('the correct writing it must not touch', () => {
// Each of these is an ordinary sentence containing a flagged substring across
// a word boundary. Without the boundary gate, every one would be corrupted —
// and corrupted silently, into text that is still made of real characters.
it('leaves two real words alone where they happen to abut', () => {
// 自己 + 经常. The substring is 己经.
expect(flagged('他自己经常做饭')).toEqual([])
// 睡觉 + 的. The substring is 觉的.
expect(flagged('睡觉的时候不要看手机')).toEqual([])
// 感觉 + 的.
expect(flagged('这是我感觉的方向')).toEqual([])
// 不知 + 到底.
expect(flagged('我不知到底该怎么办')).toEqual([])
// 因 + 位置.
expect(flagged('因位置不好我们换了座位')).toEqual([])
// 已 + 后悔.
expect(flagged('他已后悔了')).toEqual([])
})
it('leaves ordinary correct prose entirely alone', () => {
for (const good of [
'我今天早上去公园跑步了',
'他的中文说得很好',
'我已经完成了我的作业',
'现在几点了,我们再见面吧',
'我觉得这个工作很有意思',
'既然你已经知道了,就按照计划做',
]) {
expect(flagged(good), good).toEqual([])
}
})
// Where the gate costs the pack a real catch, and the trade it is making.
// 不知 is itself a word, so 我不知到他在哪里 — which really is 知到 for 知道 —
// reads to the segmenter as 不知 + 到 and is left alone. That is the gate
// preferring a missed error to a corrupted sentence, which is the whole
// premise: 我不知到底该怎么办 is the same three characters and is correct.
it('declines a real error rather than risk the sentence beside it', () => {
expect(flagged('我不知到他在哪里')).toEqual([])
expect(flagged('你知到吗')).toEqual(['知到→知道'])
})
it('says nothing about English, or about nothing', () => {
expect(flagged('I already finished my homework')).toEqual([])
expect(flagged('')).toEqual([])
})
// The direction gate. The word list is loaded only for an account learning
// Chinese, so without one this pack is silent — a writer practising English
// must never be told her own quoted Chinese is wrong.
it('is silent without a segmenter, which is how the direction gate works', () => {
expect(hanziFindings('我己经写完了', null)).toEqual([])
})
})
+149
View File
@@ -0,0 +1,149 @@
import type { MechanicsFinding } from '../../api/client'
import type { Segmenter } from '../../lib/segment'
// 错别字 — wrong-character detection, the Chinese counterpart of the spell
// checker, and a different problem from the one Hunspell solves.
//
// Chinese has no misspellings in the English sense: every character a writer can
// type is a real character, correctly formed, and an IME will not offer one that
// is not. What it *will* offer is the wrong one. Typing pinyin `yijing` and
// taking the first candidate gives 已经 or 己经 depending on the moment, and both
// are made of real characters. So the unit of error is not a malformed word but
// a **substituted character inside a correct-looking one** — which is why this
// is a rule pack over confusable pairs rather than a dictionary membership test.
//
// The discipline is Phase 22's, and the bar is the same: **precision over
// recall**. A wrong nudge costs more trust than a missed one earns, and it costs
// double here, because a learner has no way to know the tool is wrong. Two
// mechanical gates enforce it, and both are checked in the tests rather than
// asserted in prose.
// A confusable pair: `wrong` is never a word, `right` is what was meant.
//
// **Gate one — the pair must be decidable by the dictionary.** Each entry is
// admitted only if `wrong` is absent from the 188k-word list *and* `right` is
// present. That is what makes the correction a fact rather than a preference,
// and it is checked against the shipped asset in hanzi.test.ts.
//
// It is also the gate that keeps out errors everyone knows are errors. 自已 for
// 自己 is among the commonest slips in written Chinese, and 自已 is itself a
// dictionary headword — so this pack does not flag it, exactly as Phase 22's
// English pack left out `married with`. The same fate for 好象 (an older form of
// 好像, still in the dictionary), 倒底, 帐号 and 部份: all real errors by a modern
// standard, none of them decidable here.
interface Confusion {
wrong: string
right: string
// The note on the card. English, because this pack only ever runs for a writer
// whose English is the language they think in — see the direction gate below.
why: string
}
const CONFUSIONS: Confusion[] = [
// 已 / 己 / 以 — three characters that differ by one stroke and share a
// syllable. The most productive source of 错别字 there is.
{ wrong: '己经', right: '已经', why: '已经 (already) — 己 is the "self" character; the one you want is 已.' },
{ wrong: '以经', right: '已经', why: '已经 (already) — 以 is a different word; 已 is the one that means "already".' },
{ wrong: '已后', right: '以后', why: '以后 (afterwards) takes 以, not 已.' },
// 在 / 再 — same pinyin (zài), completely different jobs: one is location and
// ongoing action, the other is repetition.
{ wrong: '在见', right: '再见', why: '再见 (goodbye) — 再 is "again", which is what "see you again" needs.' },
{ wrong: '正再', right: '正在', why: '正在 (in the middle of doing) takes 在, the one about being somewhere.' },
{ wrong: '现再', right: '现在', why: '现在 (now) takes 在.' },
// 做 / 作 — both zuò, both "to do", and which one a compound takes is simply
// fixed by convention. A learner cannot reason it out, which is what makes a
// reminder worth having.
{ wrong: '工做', right: '工作', why: '工作 (work) is written with 作.' },
{ wrong: '做业', right: '作业', why: '作业 (homework) is written with 作.' },
{ wrong: '做者', right: '作者', why: '作者 (author) is written with 作.' },
{ wrong: '做文', right: '作文', why: '作文 (an essay) is written with 作.' },
{ wrong: '做用', right: '作用', why: '作用 (effect, function) is written with 作.' },
// 得 / 的 — the pair everyone knows about. Only the fixed compound is flagged:
// deciding 的 against 地 against 得 in the general case needs to know whether
// the next word is a verb or a noun, which nothing here can tell.
{ wrong: '觉的', right: '觉得', why: '觉得 (to feel, to think) ends in 得.' },
// 即 / 既 — one stroke apart, opposite meanings ("namely" against "since").
{ wrong: '即然', right: '既然', why: '既然 (since, given that) takes 既.' },
{ wrong: '既使', right: '即使', why: '即使 (even if) takes 即.' },
// The rest: ordinary IME slips where the wrong character is a homophone.
{ wrong: '因该', right: '应该', why: '应该 (should) — 因 means "because"; the word you want starts with 应.' },
{ wrong: '因位', right: '因为', why: '因为 (because) ends in 为.' },
{ wrong: '知到', right: '知道', why: '知道 (to know) ends in 道.' },
{ wrong: '安照', right: '按照', why: '按照 (according to) takes 按.' },
{ wrong: '蜜码', right: '密码', why: '密码 (password) takes 密 — 蜜 is honey.' },
{ wrong: '犹其', right: '尤其', why: '尤其 (especially) takes 尤.' },
{ wrong: '甘净', right: '干净', why: '干净 (clean) takes 干.' },
{ wrong: '什末', right: '什么', why: '什么 (what) ends in 么.' },
{ wrong: '一像', right: '一样', why: '一样 (the same) ends in 样 — 像 is "to resemble".' },
{ wrong: '必须品', right: '必需品', why: '必需品 (a necessity) takes 需. 必须 is "must", which is a different word.' },
]
// **Gate two — the characters must not already belong to two different words.**
//
// This is the gate that stops the pack from destroying correct writing, and
// without it every rule above is dangerous. 自己经常 ("oneself, often") contains
// the string 己经. 睡觉的时候 ("when sleeping") contains 觉的. 不知到底 contains 知到.
// A substring match would corrupt all three.
//
// The segmenter already knows the difference, so the test is: split the text,
// and if the two characters land in different tokens *and* either token is a
// real multi-character word, this is a word boundary and not an error. Two
// adjacent single-character tokens is what the walk produces when it has nothing
// better to offer — which is exactly what a mistyped compound looks like.
function isWordBoundary(tokens: { word: string; from: number; to: number }[], at: number): boolean {
const left = tokens.find((t) => at >= t.from && at < t.to)
const right = tokens.find((t) => at + 1 >= t.from && at + 1 < t.to)
if (!left || !right || left === right) return false
return left.word.length > 1 || right.word.length > 1
}
// hanziFindings returns the 错别字 in a piece of text, as ordinary mechanics
// findings — the same shape, the same rail, the same cards, the same accept.
//
// It needs the segmenter and does nothing without one, which is also the
// direction gate: the word list is loaded only for an account learning Chinese
// (useSegmenter), so a writer practising English can never be told her quoted
// Chinese is wrong. That is not a nicety. Petal deliberately never corrects the
// pair language — the fr and es dictionaries are chosen to hold every variety
// precisely so they cannot underline correct writing — and a Mandarin native
// does not need her own language checked by a rule pack of two dozen entries.
export function hanziFindings(text: string, segmenter: Segmenter | null): MechanicsFinding[] {
if (!segmenter || !text) return []
// One segmentation for the whole text, shared by every rule. The walk is
// linear, but running it two dozen times over a long document would not be.
const tokens = segmenter.segment(text)
const found: MechanicsFinding[] = []
for (const c of CONFUSIONS) {
let from = text.indexOf(c.wrong)
while (from !== -1) {
// The boundary test is asked at the seam the substitution sits on: the
// gap between the first two characters, which is where a mistyped
// compound and two adjacent words look different from each other.
if (!isWordBoundary(tokens, from)) {
found.push({
from,
to: from + c.wrong.length,
original: c.wrong,
replacement: c.right,
explanation: c.why,
type: 'mechanics',
})
}
from = text.indexOf(c.wrong, from + 1)
}
}
// Document order, so the rail reads down the page rather than down this file.
return found.sort((a, b) => a.from - b.from)
}
// Exported for the tests, which check every pair against the shipped word list.
// A pack whose own gate is only described in a comment is a pack whose gate can
// rot; this is how the description is made to stay true.
export const CONFUSION_PAIRS = CONFUSIONS.map((c) => ({ wrong: c.wrong, right: c.right }))
@@ -0,0 +1,79 @@
import { useEffect, useState, type RefObject } from 'react'
// How often to re-measure outside of scroll/resize events. Cards re-pack when
// suggestions arrive, expand, or get accepted — none of which fire an event we
// can hear from here, so a slow poll picks those up.
const POLL_MS = 500
// What the mascot yields to. Suggestion cards stack down into its corner; the
// History and Garden drawers cover it outright, and their footer controls sat
// under the halo. Matching on the modal role rather than each panel's own class
// means a future drawer is covered the day it's written, without a list to keep
// in sync. Anything that doesn't actually reach the corner still won't trip the
// rect test below.
const CARD = '.petal-rail-card'
const MODAL = '[role="dialog"][aria-modal="true"]'
export interface CardOverlap {
// A suggestion card reaches the mascot. It should get out of the way, but may
// still wake up when it has something to say.
cards: boolean
// A modal panel (History, Garden) covers the mascot's corner. Nothing the
// kitten wants to say is worth covering a dialog the writer opened on
// purpose, so this yields unconditionally.
modal: boolean
}
// Reports what, if anything, the mascot should yield to at the given element.
// Used to fade the corner mascot out of the way when a card or panel reaches
// into its corner, so nothing is ever hidden (or made unclickable) by the
// kitten.
export function useCardOverlap(ref: RefObject<HTMLElement | null>): CardOverlap {
const [overlap, setOverlap] = useState<CardOverlap>({ cards: false, modal: false })
useEffect(() => {
let raf = 0
const check = () => {
const el = ref.current
if (!el) return
const r = el.getBoundingClientRect()
const hits = (selector: string) => {
for (const other of document.querySelectorAll(selector)) {
const b = other.getBoundingClientRect()
if (b.left < r.right && b.right > r.left && b.top < r.bottom && b.bottom > r.top) {
return true
}
}
return false
}
const next = { cards: hits(CARD), modal: hits(MODAL) }
// Same-value object identity would re-render on every poll tick.
setOverlap((prev) =>
prev.cards === next.cards && prev.modal === next.modal ? prev : next,
)
}
const schedule = () => {
if (raf) return
raf = requestAnimationFrame(() => {
raf = 0
check()
})
}
check()
// Capture phase so scrolls inside nested scrollers (History panel, rail) count.
window.addEventListener('scroll', schedule, true)
window.addEventListener('resize', schedule)
const timer = window.setInterval(check, POLL_MS)
return () => {
window.removeEventListener('scroll', schedule, true)
window.removeEventListener('resize', schedule)
window.clearInterval(timer)
cancelAnimationFrame(raf)
}
}, [ref])
return overlap
}
+10 -1
View File
@@ -20,6 +20,12 @@ interface Props {
// The signed-in writer, when there is real auth to sign out of. Null in a
// local-dev build, where there is nothing to leave.
account: { name: string } | null
// The account's learner direction and the way to change it, passed straight
// through to the language picker in the footer — the sidebar is the drawer,
// and the drawer is the only chrome always one tap away on a phone.
direction?: string
onDirection?: (direction: string) => Promise<void>
onPair?: (lang: string, direction: string) => Promise<void>
}
// Sidebar sort orders. 'recent' keeps the server's updated_at-desc ordering.
@@ -43,6 +49,9 @@ export function DocList({
onToggleTag,
onCreateTag,
account,
direction,
onDirection,
onPair,
}: Props) {
const t = usePack()
// Active tag filter (null = show all). Cleared automatically if the tag
@@ -161,7 +170,7 @@ export function DocList({
{/* The pair Petal speaks. Unlike the rows above it this is not about any
document, and unlike sign-out it is offered whether or not there is an
account behind the session — a local-dev build still has a langpack. */}
<LanguagePicker />
<LanguagePicker direction={direction} onDirection={onDirection} onPair={onPair} />
{/* Who's writing, and the way out. Shown only when there's a real account
behind the session — a local-dev build has nobody to sign out as. */}
+100 -5
View File
@@ -15,26 +15,73 @@ import { setPackLang, shippedPacks, usePack } from '../../i18n'
// read a label that says "Portuguese" in Chinese, so the buttons say 中文 and
// Português and nothing else — the one place in Petal where bilingual copy would
// actively get in the way.
export function LanguagePicker() {
interface Props {
// The account's current direction ('learning_en' | 'learning_pair'), and the
// way to change it. Owned by App rather than here, because turning the pair
// around changes what the *editor* does — it is what loads the word list —
// and this control is only where the writer says so.
direction?: string
onDirection?: (direction: string) => Promise<void>
// Move the pair itself. Owned by App for the same reason: the answer carries
// the direction too, and the account's direction is what loads the word list.
onPair?: (lang: string, direction: string) => Promise<void>
}
export function LanguagePicker({ direction, onDirection, onPair }: Props = {}) {
const t = usePack()
const packs = shippedPacks()
const [saving, setSaving] = useState<string | null>(null)
const [failed, setFailed] = useState(false)
const [turning, setTurning] = useState(false)
const [turnFailed, setTurnFailed] = useState(false)
// Nothing to choose between — a deployment with one pack shows no picker
// rather than a single button that does nothing.
if (packs.length < 2) return null
// rather than a single button that does nothing. The direction control is
// still worth rendering in that case, so it is checked separately below.
const showPacks = packs.length >= 2
// `t.learner` is the pack's own statement that this pair can be learned
// toward, and the server keeps the matching list (auth.learnerPairs). A pack
// without it renders nothing here, which is the same failure mode as a pair
// without copy: absent rather than broken.
const learner = t.learner
if (!showPacks && !learner) return null
const turn = async (next: string) => {
if (!onDirection || next === (direction ?? 'learning_en') || turning) return
setTurning(true)
setTurnFailed(false)
try {
await onDirection(next)
} catch {
setTurnFailed(true)
} finally {
setTurning(false)
}
}
const choose = async (code: string) => {
if (code === t.code || saving) return
setSaving(code)
setFailed(false)
try {
// Name the direction alongside the pair. The server validates the two as
// one decision and refuses a learner direction for a pair it has no word
// list for, so an account that is learning Chinese cannot move to French
// by naming only the pair — that request is rejected outright, and the
// writer is left on a picker whose buttons all fail. A pair with no
// learner side can only be travelled toward English; saying so is how the
// move is actually made.
const target = packs.find((p) => p.code === code)
const next = target?.learner ? (direction ?? 'learning_en') : 'learning_en'
if (onPair) {
await onPair(code, next)
} else {
const me = await api.setPairLang(code)
// The server's answer, not the code we asked for. Everything downstream
// her dictionary, the read-aloud voice, the word lookups — follows the
// The server's answer, not the code we asked for. Everything downstream
// her dictionary, the read-aloud voice, the word lookups — follows the
// pack, so it must follow what was actually stored.
setPackLang(me.pair_lang)
}
} catch {
// A 401 has already surfaced as the sign-in overlay through the client's
// interceptor; anything else leaves her on the pair she was already on,
@@ -47,6 +94,8 @@ export function LanguagePicker() {
return (
<div className="flex flex-col gap-1 px-1">
{showPacks && (
<>
{/* Label and buttons wrap as a pair: the label is itself bilingual
("Langue · Language"), and three self-naming buttons beside it need
more than the drawer is wide in every language Petal ships. When they
@@ -89,6 +138,52 @@ export function LanguagePicker() {
{t.docs.languageFailed}
</span>
)}
</>
)}
{/* Which way round the pair is being learned. Below the language buttons
because it only makes sense once the language is settled, and rendered
at all only for a pair Petal has the learner-side data for. */}
{learner && onDirection && (
<div
className="flex flex-wrap items-center gap-x-2 gap-y-1 text-xs"
style={{ color: 'var(--color-muted)' }}
>
<span className="shrink-0 font-semibold">{learner.label}</span>
<div className="ml-auto flex shrink-0 gap-1">
{[
{ code: 'learning_en', text: learner.toEn, en: `learning English` },
{ code: 'learning_pair', text: learner.toPair, en: `learning ${t.nativeName}` },
].map((opt) => {
const active = (direction ?? 'learning_en') === opt.code
return (
<button
key={opt.code}
type="button"
onClick={() => void turn(opt.code)}
disabled={turning}
aria-pressed={active}
aria-label={`I am ${opt.en}`}
className="petal-tap-sm px-2.5 py-1 text-xs font-bold transition-colors disabled:opacity-60"
style={{
borderRadius: 'var(--radius-pill)',
background: active ? 'var(--color-accent)' : 'var(--color-surface)',
color: active ? '#fff' : 'var(--color-plum)',
boxShadow: active ? 'none' : 'var(--shadow-soft)',
}}
>
{opt.text}
</button>
)
})}
</div>
</div>
)}
{turnFailed && learner && (
<span className="text-[0.7rem]" style={{ color: 'var(--color-accent)' }}>
{learner.failed}
</span>
)}
</div>
)
}
+5 -2
View File
@@ -1,6 +1,7 @@
import { useEffect, useRef, useState } from 'react'
import { tagColorVar, type Tag, type TagColor } from '../../api/client'
import { usePack } from '../../i18n'
import { fromIME } from '../../lib/ime'
const COLORS: TagColor[] = ['rose', 'mint', 'peach', 'lavender', 'sky', 'honey']
@@ -26,7 +27,9 @@ export function TagPicker({ roster, assignedIds, onToggle, onCreate, onClose }:
if (!ref.current?.contains(e.target as Node)) onClose()
}
const onKey = (e: KeyboardEvent) => {
if (e.key === 'Escape') onClose()
// Not while an IME is open: a tag named in Chinese is composed in this
// very field, and Escape there means "wrong candidate", not "close".
if (e.key === 'Escape' && !fromIME(e)) onClose()
}
// Defer so the opening click doesn't immediately close it.
const id = setTimeout(() => document.addEventListener('pointerdown', onDown), 0)
@@ -114,7 +117,7 @@ export function TagPicker({ roster, assignedIds, onToggle, onCreate, onClose }:
value={name}
onChange={(e) => setName(e.target.value)}
onKeyDown={(e) => {
if (e.key === 'Enter') submit()
if (e.key === 'Enter' && !fromIME(e)) submit()
}}
placeholder={t.docs.newTagPlaceholder}
aria-label="New tag name"
+88 -14
View File
@@ -1,27 +1,52 @@
import { useEffect, useRef, useState } from 'react'
import { api, streamSuggestionChat, type ChatMessage } from '../../api/client'
import { usePack } from '../../i18n'
import { splitBilingual } from './bilingualReply'
import { fromIME } from '../../lib/ime'
interface Props {
suggestionId: string
// The English explanation (shown in the card body). Petal's opening bubble is
// its Simplified-Chinese translation, fetched on open — so the panel doesn't
// just repeat the same English text twice. Falls back to this on failure.
// its translation into the pair language, fetched on open — so the panel
// doesn't just repeat the same English text twice. Falls back to this on
// failure.
explanation: string
}
// CJK fallback stack — Nunito has no Chinese glyphs, and the user asks questions
// in Mandarin (spec Note #17). Applied to the bubbles specifically, not the
// serif editor body.
// CJK fallback stack — Nunito has no Chinese glyphs, and on the zh pair both
// the questions and half of every answer are in Mandarin (spec Note #17). The
// Latin pairs fall through to Nunito as before. Applied to the bubbles
// specifically, not the serif editor body.
const CHAT_FONT = "'Nunito', 'PingFang SC', 'Microsoft YaHei', 'Noto Sans CJK SC', sans-serif"
// How tall the conversation may grow (UX item 6: "room to read"). A bilingual
// three-paragraph answer in a 220px box was a scrollbar with a sentence in it.
//
// An earlier version of this took the smaller of half the viewport and the room
// left below the card, so the card could never overhang the screen. Measured, it
// gave 176px against a 442px answer — the card's own pill, diff, explanation and
// action row already spend ~290px of an 810px screen, so "fits below the word"
// and "room to read" are simply not both available.
//
// So this is the flat ceiling, and the overhang is made navigable instead —
// item 4's answer to the same conflict, and its words for it: "the answer is to
// make the overhang navigable, not to shrink what each card says". Both surfaces
// that host this panel report their reach to the editor wrapper (SuggestionCard
// via onExtent, the rail via its own measureTick), which grows the column, so a
// conversation that runs past the fold has real page under it and the Accept
// button below it can be scrolled to.
const CHAT_MAX_FRACTION = 0.5
// Below this a max-height stops being a reading area and becomes a peephole —
// the floor for a very short window, where half of it is not worth having.
const CHAT_MIN_PX = 160
// AskPetal is the mini chat panel inside an expanded SuggestionCard. The whole
// conversation lives in this component's state — nothing is persisted; closing
// the card (unmounting) clears it. Each send streams Petal's reply token-by-
// token into the latest assistant bubble.
export function AskPetal({ suggestionId, explanation }: Props) {
const t = usePack()
// Opening bubble starts empty (caret-only) and fills with the Mandarin
// Opening bubble starts empty (caret-only) and fills with the pair-language
// translation once it lands; `seeding` drives that loading caret.
const [messages, setMessages] = useState<ChatMessage[]>([{ role: 'assistant', content: '' }])
const [seeding, setSeeding] = useState(true)
@@ -36,6 +61,19 @@ export function AskPetal({ suggestionId, explanation }: Props) {
if (el) el.scrollTop = el.scrollHeight
}, [messages])
// How tall the conversation may grow. A share of the window, so a laptop and a
// large monitor both give the answer a sensible amount of themselves — and a
// window she resizes mid-conversation is answered live.
const [maxHeight, setMaxHeight] = useState(() =>
Math.max(CHAT_MIN_PX, window.innerHeight * CHAT_MAX_FRACTION),
)
useEffect(() => {
const onResize = () =>
setMaxHeight(Math.max(CHAT_MIN_PX, window.innerHeight * CHAT_MAX_FRACTION))
window.addEventListener('resize', onResize)
return () => window.removeEventListener('resize', onResize)
}, [])
// Focus the input when the panel opens. preventScroll: the card is already on
// screen as an absolutely-positioned overlay, and a default focus() would make
// the browser scroll its ancestor to "reveal" the input — jumping the document
@@ -44,7 +82,8 @@ export function AskPetal({ suggestionId, explanation }: Props) {
inputRef.current?.focus({ preventScroll: true })
}, [])
// Fetch the Chinese translation of the explanation to seed the first bubble.
// Fetch the pair-language translation of the explanation to seed the first
// bubble.
// Only replaces the seed bubble if the user hasn't started chatting yet (the
// conversation always opens with this one assistant turn). Falls back to the
// English explanation if the translation can't be fetched.
@@ -91,10 +130,10 @@ export function AskPetal({ suggestionId, explanation }: Props) {
} catch (err) {
setMessages((prev) => {
const next = prev.slice()
next[next.length - 1] = {
role: 'assistant',
content: 'Sorry, I had trouble responding just now. Please try again. 🌸',
}
// Bilingual, from the pack, and blank-line separated like a real reply —
// so the one message Petal writes without the model still renders
// through the same two-half bubble as every message with it.
next[next.length - 1] = { role: 'assistant', content: t.editor.chatFailed }
return next
})
console.error('Ask Petal chat failed:', err)
@@ -116,7 +155,7 @@ export function AskPetal({ suggestionId, explanation }: Props) {
<div
ref={scrollRef}
className="flex flex-col gap-2 overflow-y-auto pr-1"
style={{ maxHeight: 220 }}
style={{ maxHeight }}
>
{messages.map((m, i) => (
<Bubble
@@ -139,6 +178,13 @@ export function AskPetal({ suggestionId, explanation }: Props) {
ref={inputRef}
value={input}
onChange={(e) => setInput(e.target.value)}
// She asks Petal in Mandarin, so the Enter that commits an IME
// candidate lands in this field constantly. Most browsers already
// withhold implicit form submission during a composition; the ones
// that don't would send her half-typed question. Cheap to be certain.
onKeyDown={(e) => {
if (e.key === 'Enter' && fromIME(e)) e.preventDefault()
}}
placeholder={t.editor.askPlaceholder}
className="min-w-0 flex-1 rounded-full px-3 py-1.5 text-xs focus:outline-none"
style={{
@@ -164,6 +210,19 @@ export function AskPetal({ suggestionId, explanation }: Props) {
// Bubble renders one chat turn: Petal rose-tinted and left-aligned, the user
// lavender and right-aligned. A trailing caret marks the actively streaming
// reply until its first token lands.
//
// Petal's turns are bilingual (see bilingualReply.ts) and are laid out the way
// the companion lays out its own two lines: the pair language first and plainly
// readable, the English beneath it in the muted tone. That order is the pack's
// order everywhere else in the UI, and it holds whichever direction the writer
// is learning in — the muted half is the one they can already read, and which
// half that is isn't Petal's to decide. The writer's own turns are their own
// words in whichever language they typed them, so they are never split.
//
// Petal's bubble also takes the full width the card offers rather than the 85%
// a chat normally reserves to show who is talking — the alignment and the
// tint already say that, and two languages in a 4/5-width column wraps a
// sentence-length answer into a paragraph-shaped one.
function Bubble({
role,
content,
@@ -174,10 +233,11 @@ function Bubble({
streaming: boolean
}) {
const isPetal = role === 'assistant'
const reply = isPetal ? splitBilingual(content) : null
return (
<div className={`flex ${isPetal ? 'justify-start' : 'justify-end'}`}>
<div
className="max-w-[85%] rounded-2xl px-3 py-1.5 text-xs leading-snug"
className={`${isPetal ? 'w-full' : 'max-w-[85%]'} rounded-2xl px-3 py-2 leading-snug`}
style={{
background: isPetal ? 'var(--color-surface-alt)' : 'var(--color-lavender)',
color: 'var(--color-plum)',
@@ -185,7 +245,21 @@ function Bubble({
whiteSpace: 'pre-wrap',
}}
>
{content}
{reply ? (
<>
<span className="text-[0.8rem]">{reply.native}</span>
{reply.en !== '' && (
<span
className="mt-1.5 block text-xs"
style={{ color: 'var(--color-muted)' }}
>
{reply.en}
</span>
)}
</>
) : (
<span className="text-xs">{content}</span>
)}
{streaming && content === '' && (
<span className="petal-chat-caret" aria-hidden>
@@ -0,0 +1,235 @@
import { describe, it, expect } from 'vitest'
import { EditorState, TextSelection } from '@tiptap/pm/state'
import type { Transaction } from '@tiptap/pm/state'
import { Schema } from '@tiptap/pm/model'
import { compositionKey, compositionPlugin, isComposing, holdRedraw } from './Composition'
import { suggestionPlugin, suggestionPluginKey, setSuggestions } from './SuggestionHighlight'
import { spellPlugin, spellPluginKey, setSpellChecker } from './SpellCheck'
import { searchPlugin, searchPluginKey, setSearch } from './SearchHighlight'
import type { Suggestion } from '../../api/client'
import type { SpellChecker } from '../../hooks/useSpellChecker'
// These tests are about one moment: she is typing 公园 with a pinyin IME, so the
// document briefly contains "gongyuan" and a candidate window sits over it. Every
// decoration layer wants to recompute, and recomputing rewrites the DOM around
// the node the browser is composing in — which is what eats half-typed input.
//
// Nothing here needs a real EditorView: composition is tracked in plugin state
// by the compositionstart/compositionend handlers, so a plain EditorState with
// the same plugins reproduces exactly the decisions the layers make.
const schema = new Schema({
nodes: {
doc: { content: 'block+' },
paragraph: { group: 'block', content: 'inline*', toDOM: () => ['p', 0] },
text: { group: 'inline' },
},
})
const doc = (text: string) =>
schema.node('doc', null, [schema.node('paragraph', null, text ? [schema.text(text)] : [])])
// A dictionary that knows ordinary English and nothing else — so the pinyin run
// an IME leaves in the document mid-composition is a misspelling to it, which is
// precisely the risk this guard exists for.
const english: SpellChecker = {
correct: (w) => ['the', 'park', 'went', 'to', 'today'].includes(w.toLowerCase()),
suggest: () => [],
extendedAlphabet: false,
}
const suggestion = (original: string, replacement: string): Suggestion => ({
id: `s-${original}`,
doc_id: 'd',
from_pos: 0,
to_pos: 0,
original,
replacement,
explanation: '',
type: 'grammar',
status: 'pending',
source: 'llm',
created_at: new Date().toISOString(),
})
function harness(text: string) {
let state = EditorState.create({
schema,
doc: doc(text),
plugins: [compositionPlugin({ onEnd: null }), suggestionPlugin(), spellPlugin(), searchPlugin()],
})
const api = {
get state() {
return state
},
tr: (f: (tr: Transaction) => Transaction) => {
state = state.apply(f(state.tr))
},
dispatch: (tr: Transaction) => {
state = state.apply(tr)
},
// The two ends of a composition, as the DOM handlers dispatch them.
startComposing: () => api.tr((tr) => tr.setMeta(compositionKey, true)),
endComposing: () => api.tr((tr) => tr.setMeta(compositionKey, false)),
// Typing, whether by keystroke or by an IME writing into the document.
type: (at: number, text: string) =>
api.tr((tr) => tr.insertText(text, at).setSelection(TextSelection.create(tr.doc, at + text.length))),
// Replace a span, the way committing an IME candidate does.
commit: (from: number, to: number, text: string) => api.tr((tr) => tr.insertText(text, from, to)),
spans: (key: typeof suggestionPluginKey | typeof spellPluginKey | typeof searchPluginKey) => {
const deco = (key.getState(state) as { decorations: import('@tiptap/pm/view').DecorationSet }).decorations
return deco.find().map((d) => [d.from, d.to] as const)
},
}
return api
}
describe('composition tracking', () => {
it('is off until a composition starts, and off again once it ends', () => {
const h = harness('I went to the ')
expect(isComposing(h.state)).toBe(false)
h.startComposing()
expect(isComposing(h.state)).toBe(true)
h.endComposing()
expect(isComposing(h.state)).toBe(false)
})
it('releases the redraw on the very transaction that ends the composition', () => {
const h = harness('hello')
h.startComposing()
const before = h.state
expect(holdRedraw(before.tr, before)).toBe(true)
// The end transaction is dispatched while composing is still true; if it
// held its own redraw like any other, nothing would ever release it.
expect(holdRedraw(before.tr.setMeta(compositionKey, false), before)).toBe(false)
})
})
describe('spell underlines during composition', () => {
it('does not underline the pinyin she is part-way through converting', () => {
const h = harness('I went to the ')
h.dispatch(h.state.tr.setMeta(spellPluginKey, english))
expect(h.spans(spellPluginKey)).toEqual([])
h.startComposing()
// The IME writes its buffer into the document one letter at a time. The
// caret sits inside the run, so the caret exemption would cover "gongyuan"
// on its own — but not a second word, and not once she moves back to fix a
// syllable. The guard is what makes that irrelevant.
h.type(15, 'gong')
h.type(19, 'yuan')
expect(h.spans(spellPluginKey)).toEqual([])
// And the caret has moved away, which normally forces a rebuild.
h.tr((tr) => tr.setSelection(TextSelection.create(tr.doc, 1)))
expect(h.spans(spellPluginKey)).toEqual([])
})
it('re-checks the moment the candidate is committed', () => {
const h = harness('I went to the ')
h.dispatch(h.state.tr.setMeta(spellPluginKey, english))
h.startComposing()
h.type(15, 'gongyuan')
h.commit(15, 23, '公园') // she picks 公园; the pinyin is gone
h.endComposing()
// Nothing to flag: the pinyin never existed by the time anyone looked, and
// CJK is not tokenized at all.
expect(h.spans(spellPluginKey)).toEqual([])
// A real misspelling typed afterwards still underlines, so the layer is
// released rather than switched off. (The caret moves off it first: a word
// under the cursor is exempt, mid-typing, IME or no IME.)
h.type(17, ' parc')
h.tr((tr) => tr.setSelection(TextSelection.create(tr.doc, 1)))
expect(h.spans(spellPluginKey).length).toBe(1)
})
it('underlines the same text immediately when no IME is involved', () => {
const h = harness('I went to the ')
h.dispatch(h.state.tr.setMeta(spellPluginKey, english))
h.type(15, 'gongyuan')
h.tr((tr) => tr.setSelection(TextSelection.create(tr.doc, 1)))
expect(h.spans(spellPluginKey).length).toBe(1)
})
})
describe('suggestion highlights during composition', () => {
it('carries existing highlights along with the text instead of re-anchoring', () => {
const h = harness('I went to the park today')
setSuggestions(h.state, h.dispatch, [suggestion('went to', 'go to')])
expect(h.spans(suggestionPluginKey)).toEqual([[3, 10]])
h.startComposing()
h.type(1, 'x') // insert before the highlight: it has to move with the text
expect(h.spans(suggestionPluginKey)).toEqual([[4, 11]])
})
it('holds a freshly arrived suggestion list until the composition ends', () => {
const h = harness('I went to the park today')
h.startComposing()
setSuggestions(h.state, h.dispatch, [suggestion('the park', 'a park')])
// The list is stored, but the page is not repainted under the IME.
expect(h.spans(suggestionPluginKey)).toEqual([])
h.endComposing()
expect(h.spans(suggestionPluginKey)).toEqual([[11, 19]])
})
it('re-anchors against the committed text, not the pinyin it replaced', () => {
const h = harness('I went to ')
setSuggestions(h.state, h.dispatch, [suggestion('公园', '花园')])
expect(h.spans(suggestionPluginKey)).toEqual([]) // not there yet
h.startComposing()
h.type(11, 'gongyuan')
h.commit(11, 19, '公园')
h.endComposing()
expect(h.spans(suggestionPluginKey)).toEqual([[11, 13]])
})
})
describe('find-and-replace highlights during composition', () => {
it('holds the match set, then refreshes it against the committed text', () => {
const h = harness('公园 and 公园')
setSearch(h.state, h.dispatch, '公园', false)
expect(h.spans(searchPluginKey).length).toBe(2)
h.startComposing()
h.type(10, ' gongyuan') // at the end of the text, where the caret is
expect(h.spans(searchPluginKey).length).toBe(2) // still two, not three
h.commit(11, 19, '公园')
h.endComposing()
expect(h.spans(searchPluginKey).length).toBe(3)
})
it('closing the bar clears immediately — a composition never holds a removal', () => {
const h = harness('公园 and 公园')
setSearch(h.state, h.dispatch, '公园', false)
h.startComposing()
h.dispatch(h.state.tr.setMeta(searchPluginKey, { kind: 'clear' }))
expect(h.spans(searchPluginKey)).toEqual([])
})
})
describe('the layers are only paused, never left stale', () => {
it('rebuilds even if the composition ends on a transaction of its own', () => {
// The end signal is dispatched on a timer, after ProseMirror has flushed the
// composition's last document change — so the releasing transaction usually
// carries no document change at all. That must still be enough.
const h = harness('I went to the ')
h.dispatch(h.state.tr.setMeta(spellPluginKey, english))
h.startComposing()
h.type(15, 'parc')
expect(h.spans(spellPluginKey)).toEqual([])
h.tr((tr) => tr.setSelection(TextSelection.create(tr.doc, 1)))
h.endComposing() // no doc change, no selection change
expect(h.spans(spellPluginKey).length).toBe(1)
})
it('a checker arriving mid-composition is applied once it ends', () => {
const h = harness('公园 parc')
h.startComposing()
setSpellChecker(h.state, h.dispatch, english)
expect(h.spans(spellPluginKey)).toEqual([])
h.tr((tr) => tr.setSelection(TextSelection.create(tr.doc, 1)))
h.endComposing()
expect(h.spans(spellPluginKey).length).toBe(1)
})
})
+112
View File
@@ -0,0 +1,112 @@
import { Extension } from '@tiptap/core'
import { Plugin, PluginKey } from '@tiptap/pm/state'
import type { EditorState, Transaction } from '@tiptap/pm/state'
import type { EditorView } from '@tiptap/pm/view'
// Composition tracks whether an IME composition is in flight, and is the one
// place the rest of the editor asks.
//
// Why it exists: typing Chinese (or Japanese, or Korean) does not produce
// characters a keystroke at a time. The IME opens a *composition* — the pinyin
// she types goes into the document as it is typed, a candidate window sits over
// it, and only when she picks a candidate is the run replaced with hanzi.
// Petal's three decoration layers (SuggestionHighlight, SpellCheck,
// SearchHighlight) all recompute from the live document on every change, so
// mid-composition they would recompute over half-typed pinyin — and rebuilding
// decorations means rewriting the DOM around the node the IME is composing in.
// That is the classic bug that eats half-typed input: the composition is
// abandoned by the browser and the letters vanish or double.
//
// The fix is to hold the redraws, not to skip them. Decorations that are due
// while a composition is in flight are kept (mapped through the transaction, so
// they follow the text that moved) and rebuilt the moment the composition ends.
// Nothing is lost — the pause is measured in the length of one word.
//
// Input rules need no guard here: Tiptap's own input-rule plugin already returns
// early while `view.composing` is true, which matters because pinyin uses an
// apostrophe as a syllable separator (xi'an → 西安) and Typography.ts rewrites
// every ' into a curly .
export const compositionKey = new PluginKey<boolean>('petalComposition')
// isComposing answers "was an IME composition in flight as of this state?".
// Decoration plugins ask it of the state *before* the transaction they are
// applying, which is what makes the answer independent of plugin ordering: the
// flag was set by an earlier transaction (compositionstart), not by this one.
export function isComposing(state: EditorState): boolean {
return compositionKey.getState(state) === true
}
// holdRedraw is the question every decoration layer asks in its `apply`: should
// this rebuild wait? Yes while composing — except on the transaction that ends
// the composition, which is precisely the one that releases the held redraws.
export function holdRedraw(tr: Transaction, stateBefore: EditorState): boolean {
if (tr.getMeta(compositionKey) === false) return false
return isComposing(stateBefore)
}
function setComposing(view: EditorView, composing: boolean) {
if (compositionKey.getState(view.state) === composing) return
view.dispatch(view.state.tr.setMeta(compositionKey, composing))
}
export interface CompositionOptions {
// Called once after a composition has ended and the document has settled.
// EditorCore uses it to re-report the committed text, since the analysis
// passes were told to ignore everything typed while composing.
onEnd: (() => void) | null
}
export function compositionPlugin(options: CompositionOptions): Plugin<boolean> {
return new Plugin<boolean>({
key: compositionKey,
state: {
init: () => false,
apply(tr, value) {
const meta = tr.getMeta(compositionKey)
return typeof meta === 'boolean' ? meta : value
},
},
props: {
handleDOMEvents: {
compositionstart: (view) => {
setComposing(view, true)
return false
},
// A custom handleDOMEvents handler runs *before* ProseMirror's own, and
// ProseMirror's compositionend queues the composition's final DOM
// changes as a microtask. Ending on a macrotask puts us after both, so
// the rebuild we release sees the committed hanzi rather than the pinyin
// it replaced. (If a transaction from that flush arrives first it
// rebuilds anyway — by then `composing` is false. Both orders land.)
compositionend: (view) => {
setTimeout(() => {
if (view.isDestroyed) return
setComposing(view, false)
options.onEnd?.()
}, 0)
return false
},
// Clicking away mid-candidate abandons the composition without a
// compositionend in some browsers. Without this the layers would stay
// held — silently, and until she typed again.
blur: (view) => {
setComposing(view, false)
return false
},
},
},
})
}
export const Composition = Extension.create<CompositionOptions>({
name: 'composition',
addOptions() {
return { onEnd: null }
},
addProseMirrorPlugins() {
return [compositionPlugin(this.options)]
},
})
+315 -43
View File
@@ -15,10 +15,11 @@ import TableHeader from '@tiptap/extension-table-header'
import TableCell from '@tiptap/extension-table-cell'
import { FontSize } from './FontSize'
import type { EditorView } from '@tiptap/pm/view'
import { useCallback, useEffect, useRef, useState } from 'react'
import { useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState } from 'react'
import { Toolbar } from '../Toolbar/Toolbar'
import { SuggestionCard } from './SuggestionCard'
import { SuggestionRail, type RailItem } from './SuggestionRail'
import { railFitsBeside } from './railFit'
import { SuggestionHighlight, setSuggestions, setActiveSuggestion, findRange } from './SuggestionHighlight'
import { SpellCheck, setSpellChecker, wordAt } from './SpellCheck'
import { MisspellCard } from './MisspellCard'
@@ -28,16 +29,35 @@ import { SelectionBubble } from './SelectionBubble'
import { SearchHighlight } from './SearchHighlight'
import { FindReplace } from './FindReplace'
import { Typography } from './Typography'
import { Composition } from './Composition'
import { RewritePreview, type RewriteStatus } from './RewritePreview'
import { api, type Suggestion, type WordInfo } from '../../api/client'
import { planBatch } from './acceptBatch'
import { api, type Suggestion, type SuggestionType, type WordInfo } from '../../api/client'
import { speak, speechSupported } from '../../audio/speech'
import type { SpellChecker } from '../../hooks/useSpellChecker'
import type { Segmenter } from '../../lib/segment'
import { hanziWordAt, hanziToWordInfo, hanziPinyin } from './hanziWord'
import { usePack } from '../../i18n'
// Breathing room left below the last suggestion card when the rail's stack is what
// defines the column's height, so the bottom card doesn't sit flush on the edge.
const RAIL_TAIL = 24
export interface EditorChange {
content: string // Tiptap JSON, stringified
content_text: string // flattened plain text for the LLM
word_count: number
// True while an IME composition is in flight: this text contains the pinyin
// she is part-way through converting, not the sentence she is writing.
//
// The save is deliberately NOT gated on it — a tablet keyboard can hold one
// composition open for a whole sentence, and Petal never makes writing wait
// for anything. Saving an intermediate state costs nothing: the next change
// supersedes it, and one always arrives (this component emits a final change
// once the composition commits). What it gates is *analysis* — asking the
// rule pack or the model to read half-typed pinyin can only produce advice
// about text that is about to stop existing.
composing: boolean
}
interface Props {
@@ -49,6 +69,9 @@ interface Props {
// call + state removal (accept's text replacement happens here in the editor).
suggestions: Suggestion[]
onAccept: (s: Suggestion) => void
// A whole category accepted at once. The text replacement happens here, in one
// transaction; the parent files each row and cheers once for the batch.
onAcceptMany: (list: Suggestion[]) => void
onDismiss: (s: Suggestion) => void
// Triggers the whole-document voice-consistency pass; `voicing` is true while
// it runs (drives the toolbar button's loading state).
@@ -64,6 +87,11 @@ interface Props {
// to the personal dictionary is bubbled up so it persists app-wide.
spellChecker: SpellChecker | null
onAddWord: (word: string) => void
// The Chinese word list, non-null only for a writer learning the pair
// language (users.direction = 'learning_pair'). Its presence is what turns on
// every Chinese-side behaviour here: hanzi stops being text the editor steps
// over and becomes words it can point at.
segmenter: Segmenter | null
}
interface MisspellState {
@@ -86,6 +114,12 @@ interface WordInfoState {
left: number
loading: boolean
info: WordInfo | null
// The word's own pinyin, for a Chinese lookup. Kept beside `info` rather than
// inside it because WordInfo is the English dictionary's shape and `phonetic`
// there means IPA — printing pinyin between the slashes that say "this is
// IPA" would be a small lie in the one place a learner is looking for the
// truth about pronunciation.
pinyin: string
// Garden state: the captured word's id (null until the auto-capture returns or
// after it's removed) and whether it's currently in the garden.
vocabId: string | null
@@ -179,6 +213,10 @@ interface GlossState {
gloss: string
// The other reading, when the token is a word in her language too.
reverse?: string
// A line shown *above* the meaning rather than below it: pinyin, for a
// Chinese word. Above because it is read first — the meaning of 公园 may
// already be clear to someone who cannot yet say it.
lead?: string
from: number
to: number
top: number
@@ -216,6 +254,7 @@ export function EditorCore({
onChange,
suggestions,
onAccept,
onAcceptMany,
onDismiss,
onVoiceCheck,
voicing,
@@ -224,11 +263,15 @@ export function EditorCore({
onFocusMode,
spellChecker,
onAddWord,
segmenter,
}: Props) {
// Her pair's copy — the hover tip labels the second reading with the language's
// own name, so it says "português" rather than "pt-PT".
const pack = usePack()
const wrapperRef = useRef<HTMLDivElement>(null)
// The text column itself, measured separately from its wrapper: the wrapper is
// grown to cover the card stack, so only this reports the height of the prose.
const contentRef = useRef<HTMLDivElement>(null)
const [hover, setHover] = useState<HoverState | null>(null)
// The open spelling popover (click a red-underlined word), or null.
const [misspell, setMisspell] = useState<MisspellState | null>(null)
@@ -266,12 +309,46 @@ export function EditorCore({
// `activeId` is the suggestion currently emphasized (hovered text or card).
const [railItems, setRailItems] = useState<RailItem[]>([])
const [railEnabled, setRailEnabled] = useState(false)
// How far the resolved card stack reaches below the wrapper's top, reported by
// the rail. Cards are absolutely positioned and so contribute no layout height:
// without this the column below the last line of text isn't scrollable and any
// card that lands there is unreachable, not merely far from its sentence.
const [railExtent, setRailExtent] = useState(0)
// The same report from the anchored card, which is absolutely positioned for
// the same reason and so has the same problem: an open Ask Petal conversation
// can reach well past the last line of a short document, and its Accept button
// goes with it. 0 whenever no card is open.
const [cardExtent, setCardExtent] = useState(0)
// How far down the column has to reach to cover its floating surfaces. Both
// reach past the prose for the same reason and are answered the same way, so
// they resolve to one number: whichever is lower wins, and 0 means the text
// alone decides the height.
//
// The rail's extent is conditional on the rail being mounted — a stale measure
// from a rail that has since been dismissed would leave a document padded with
// blank scroll. The card's is not: it reports 0 as it unmounts.
const overhang = Math.max(
railEnabled && railExtent > 0 ? railExtent + RAIL_TAIL : 0,
cardExtent > 0 ? cardExtent + RAIL_TAIL : 0,
)
// Sticky offset for the text column, or null when it should sit in normal flow.
// Set only while the stack overhangs the text: scrolling down to reach the lower
// cards would otherwise carry every sentence off the top of the screen.
const [stickTop, setStickTop] = useState<number | null>(null)
const [railExpandedId, setRailExpandedId] = useState<string | null>(null)
const [activeId, setActiveId] = useState<string | null>(null)
// A stable handle to the latest recompute so the editor's onUpdate (captured
// once at construction) can trigger a re-measure without stale closures.
const recomputeRailRef = useRef<() => void>(() => {})
// Re-report the document once an IME composition commits. Everything typed
// while composing was reported with `composing: true`, so the analysis passes
// ignored it; without this nudge the committed sentence would wait for the
// next keystroke to be looked at. Held in a ref because the extension list is
// built once, at construction.
const emitCommittedRef = useRef<() => void>(() => {})
const editor = useEditor({
extensions: [
StarterKit,
@@ -289,6 +366,11 @@ export function EditorCore({
TextAlign.configure({ types: ['heading', 'paragraph'] }),
Placeholder.configure({ placeholder: 'Start writing…' }),
CharacterCount,
// First in the list so its state is settled before the layers that read
// it — not that they depend on the ordering (they read the state as of
// the transaction before), but the one that answers the question should
// come before the ones that ask it.
Composition.configure({ onEnd: () => emitCommittedRef.current() }),
SuggestionHighlight,
SpellCheck,
SearchHighlight,
@@ -333,6 +415,7 @@ export function EditorCore({
content: JSON.stringify(editor.getJSON()),
content_text: editor.getText(),
word_count: editor.storage.characterCount.words(),
composing: editor.view.composing,
})
// Edits reflow the text, so the rail anchors need re-measuring.
recomputeRailRef.current()
@@ -359,6 +442,26 @@ export function EditorCore({
},
})
// The composition-end nudge. Same payload as onUpdate's, with `composing`
// false by construction — this runs after the composition has ended and its
// final changes have been flushed, so the text here is the committed one.
//
// Written in an effect rather than during render, like recomputeRailRef
// below: a render React throws away must not be the one that leaves its
// closure behind for a DOM event to call later.
useEffect(() => {
emitCommittedRef.current = () => {
if (!editor) return
onChange({
content: JSON.stringify(editor.getJSON()),
content_text: editor.getText(),
word_count: editor.storage.characterCount.words(),
composing: false,
})
recomputeRailRef.current()
}
}, [editor, onChange])
// When the selected document changes, swap in its content without emitting an
// update (false) so loading a doc doesn't trigger a spurious save.
useEffect(() => {
@@ -381,6 +484,26 @@ export function EditorCore({
// popover would offer a definition of "cora".
const wordAlphabet = spellChecker?.extendedAlphabet ?? false
// "The word under here", for a document that may hold two writing systems at
// once — which every document in this pair does, because a learner's Chinese
// practice is full of English and her English is full of quoted Chinese.
//
// Chinese is tried first and Latin second, and the order costs nothing to get
// right: the two can never both answer, because a Han character is not a Latin
// letter and neither tokenizer will cross into the other's run. `hanzi` rides
// along because the two answers go to different dictionaries — the same
// string is a word in exactly one of them.
const resolveWord = useCallback(
(pos: number): { from: number; to: number; word: string; hanzi: boolean } | null => {
if (!editor) return null
const han = hanziWordAt(editor.state.doc, pos, segmenter)
if (han) return { ...han, hanzi: true }
const latin = wordAt(editor.state.doc, pos, wordAlphabet)
return latin ? { ...latin, hanzi: false } : null
},
[editor, segmenter, wordAlphabet],
)
// Push the spell checker into its decoration plugin once the dictionary loads
// (and again whenever the personal dictionary changes its identity).
useEffect(() => {
@@ -409,9 +532,7 @@ export function EditorCore({
const wrapper = wrapperRef.current
if (!wrapper) return
const wrapRect = wrapper.getBoundingClientRect()
// Need room for the 300px column + its 32px gutter (see .petal-rail), plus
// a little breathing space to the viewport edge.
setRailEnabled(window.innerWidth - wrapRect.right >= 348)
setRailEnabled(railFitsBeside(window.innerWidth, wrapRect.right))
const seen = new Set<string>()
const items: RailItem[] = []
wrapper.querySelectorAll<HTMLElement>('.petal-suggestion[data-suggestion-id]').forEach((el) => {
@@ -432,11 +553,24 @@ export function EditorCore({
// Re-anchor when the suggestion set changes (after the decorations repaint),
// and keep the rail in sync with viewport/editor width changes (room + reflow).
//
// The scrollport is observed as well as the wrapper, and it is not redundant:
// the wrapper is a fixed 720px column, so entering or leaving distraction-free
// mode *moves* it (the pane re-centres) without ever changing its size. A
// ResizeObserver on the wrapper alone reports nothing, no window resize fires,
// and `railEnabled` keeps whatever value it had — leaving the 300px rail
// rendered into the 266px margin a restored sidebar leaves behind, cards
// clipped mid-sentence and the page scrolling sideways. The scrollport spans
// the pane, so it resizes whenever the chrome around the editor does.
useEffect(() => {
recomputeRail()
const wrapper = wrapperRef.current
const port = wrapper?.closest('.petal-scrollport')
const ro = wrapper ? new ResizeObserver(() => recomputeRail()) : null
if (wrapper && ro) ro.observe(wrapper)
if (wrapper && ro) {
ro.observe(wrapper)
if (port) ro.observe(port)
}
window.addEventListener('resize', recomputeRail)
return () => {
ro?.disconnect()
@@ -444,6 +578,42 @@ export function EditorCore({
}
}, [recomputeRail])
// The rail only reports its extent while it's mounted, so clear it when the last
// card goes (accepted the lot, or the window narrowed past the rail's threshold)
// — otherwise the column keeps the height of a stack that no longer exists.
useEffect(() => {
if (!railEnabled || railItems.length === 0) setRailExtent(0)
}, [railEnabled, railItems.length])
// Decide whether the text column has to be pinned. The rail's cards hang off an
// absolutely-positioned column, so when several suggestions share one short
// paragraph the stack runs far past the last line of text. Growing the wrapper to
// `railExtent` makes that space scrollable (item 4: the lower cards were simply
// unreachable); pinning the prose inside it means scrolling down to read those
// cards keeps the sentences on screen instead of scrolling them away.
//
// The offset is `min(0, port - content)`: prose shorter than the viewport sticks
// at the top, taller prose sticks by its *bottom* edge, so its last lines — the
// ones the overhanging cards flag — stay visible rather than the first.
useLayoutEffect(() => {
const wrapper = wrapperRef.current
const content = contentRef.current
if (!wrapper || !content || !railEnabled || railExtent <= 0) {
setStickTop(null)
return
}
const contentH = content.offsetHeight
// Only pin when the stack actually overhangs the prose; a rail that fits
// beside its text needs nothing, and pinning it would be a change for free.
if (railExtent <= contentH) {
setStickTop(null)
return
}
const port = wrapper.closest('.petal-scrollport')
const portH = port ? port.clientHeight : window.innerHeight
setStickTop(Math.min(0, portH - contentH - RAIL_TAIL))
}, [railEnabled, railExtent, railItems])
// Emphasize the flagged text for the active suggestion, mirroring the rail
// card ↔ text link both ways. Driven through the decoration plugin (not an
// imperative DOM class) so it survives the repaints that fire on every edit.
@@ -538,12 +708,15 @@ export function EditorCore({
(e: React.MouseEvent) => {
if (!(e.target as HTMLElement).closest('.petal-suggestion')) return
if (railEnabled) {
setActiveId(null)
// A click-opened card outlives the pointer (it closes on a click away),
// so its rail card keeps the glow — otherwise the open card and the
// margin stop agreeing about which suggestion is being read.
if (!hover) setActiveId(null)
return
}
scheduleClose()
},
[scheduleClose, railEnabled],
[scheduleClose, railEnabled, hover],
)
const keepOpen = useCallback(() => clearTimeout(closeTimer.current), [])
@@ -553,18 +726,28 @@ export function EditorCore({
// anchored to the highlight itself (captured before the replacement removes it),
// so it fires in the right spot whether the accept came from the hover card or
// the margin rail.
const handleAccept = useCallback(
(s: Suggestion) => {
// Where a suggestion's highlight sits, in wrapper coordinates — captured before
// the replacement removes it, so the confetti fires over the words that changed.
const burstAt = useCallback((id: string): { top: number; left: number } | null => {
const wrapper = wrapperRef.current
const el = wrapper?.querySelector(
`.petal-suggestion[data-suggestion-id="${CSS.escape(s.id)}"]`,
`.petal-suggestion[data-suggestion-id="${CSS.escape(id)}"]`,
) as HTMLElement | null
let burst: { top: number; left: number } | null = null
if (wrapper && el) {
if (!wrapper || !el) return null
const wrapRect = wrapper.getBoundingClientRect()
const elRect = el.getBoundingClientRect()
burst = { top: elRect.top - wrapRect.top, left: elRect.right - wrapRect.left }
}
return { top: elRect.top - wrapRect.top, left: elRect.right - wrapRect.left }
}, [])
const showConfetti = useCallback((burst: { top: number; left: number }) => {
setConfetti(burst)
clearTimeout(confettiTimer.current)
confettiTimer.current = setTimeout(() => setConfetti(null), 720)
}, [])
const handleAccept = useCallback(
(s: Suggestion) => {
let burst = burstAt(s.id)
if (editor && s.replacement.trim() !== '') {
const range = findRange(editor.state.doc, s.original)
if (range) {
@@ -573,16 +756,57 @@ export function EditorCore({
}
// Fall back to the hover card's position if the highlight wasn't found.
if (!burst && hover) burst = { top: hover.top, left: hover.left + 16 }
if (burst) {
setConfetti(burst)
clearTimeout(confettiTimer.current)
confettiTimer.current = setTimeout(() => setConfetti(null), 720)
}
if (burst) showConfetti(burst)
closeCard()
setRailExpandedId(null)
onAccept(s)
},
[editor, onAccept, closeCard, hover],
[editor, onAccept, closeCard, hover, burstAt, showConfetti],
)
// How many pending cards of each type could be accepted in one go. A card
// offers "Accept all" only when it has company, so a lone Grammar card doesn't
// grow a second button saying the same thing as the first.
const batchCounts = useMemo(() => {
const counts: Partial<Record<SuggestionType, number>> = {}
for (const s of suggestions) {
if (s.replacement.trim() === '') continue // awareness-only: nothing to accept
counts[s.type] = (counts[s.type] ?? 0) + 1
}
return counts
}, [suggestions])
// Accept a whole category at once. Five tense fixes were five clicks, five
// confetti bursts and five separate undo steps; this is one of each. The single
// undo step is the reason every replacement goes into ONE chain: Tiptap applies
// a chain as a single transaction, and prosemirror-history undoes it as a
// single event, so Ctrl+Z takes back the batch rather than unpicking it.
const handleAcceptAll = useCallback(
(type: SuggestionType) => {
if (!editor) return
const family = suggestions.filter((s) => s.type === type)
const plan = planBatch(family, (original) => findRange(editor.state.doc, original))
const settled = [...plan.steps.map((step) => step.suggestion), ...plan.missing]
if (settled.length === 0) return
// The topmost span that's about to change — the last step, since steps run
// bottom-up. Read before the edit, while the highlights still exist.
const first = plan.steps[plan.steps.length - 1]
const burst = first ? burstAt(first.suggestion.id) : null
if (plan.steps.length > 0) {
let chain = editor.chain().focus()
for (const step of plan.steps) {
chain = chain.insertContentAt({ from: step.from, to: step.to }, step.suggestion.replacement)
}
chain.run()
}
if (burst) showConfetti(burst)
closeCard()
setRailExpandedId(null)
onAcceptMany(settled)
},
[editor, suggestions, onAcceptMany, closeCard, burstAt, showConfetti],
)
const handleDismiss = useCallback(
@@ -633,15 +857,15 @@ export function EditorCore({
if (suggestionEl) {
const id = suggestionEl.getAttribute('data-suggestion-id')
if (id) {
// With the rail open, a tap emphasizes and expands its margin card
// instead of opening a floating one.
if (railEnabled) {
// Clicking a highlight always opens the card at the word. Hover still
// defers to the rail (see handleMouseOver) — an unbidden floating card
// beside a margin card that already says the same thing is noise. But a
// click is her asking to deal with *this* word, and answering it 650px
// away in the periphery is the gesture item 7 is about. The rail card
// glows rather than expanding, so the suggestion is never open twice.
setActiveId(id)
setRailExpandedId(id)
} else {
openCardFor(id, suggestionEl)
}
}
return
}
if (!editor) return
@@ -692,7 +916,7 @@ export function EditorCore({
const openWordLookup = useCallback(
(pos: number) => {
if (!editor) return
const range = wordAt(editor.state.doc, pos, wordAlphabet)
const range = resolveWord(pos)
if (!range) return
const wrapper = wrapperRef.current
if (!wrapper) return
@@ -708,12 +932,18 @@ export function EditorCore({
closeCard()
setMisspell(null)
const token = ++wordReqRef.current
setWordInfo({ word: range.word, from: range.from, to: range.to, top, left, loading: true, info: null, vocabId: null, saved: false })
setWordInfo({ word: range.word, from: range.from, to: range.to, top, left, loading: true, info: null, pinyin: '', vocabId: null, saved: false })
// The sentence the word sits in, for review context in the garden.
const example = exampleAt(range.from)
api
.lookupWord(range.word)
.then((info) => {
// Two dictionaries, one card. The Chinese lookup answers in English and
// the English one answers in her language; which is wanted follows from
// which script the word is written in, so nothing here has to consult the
// account's direction a second time.
const lookup: Promise<{ info: WordInfo; pinyin: string }> = range.hanzi
? api.hanziWord(range.word).then((h) => ({ info: hanziToWordInfo(h), pinyin: hanziPinyin(h) }))
: api.lookupWord(range.word).then((info) => ({ info, pinyin: '' }))
lookup
.then(({ info, pinyin }) => {
if (token !== wordReqRef.current) return
// Auto-capture into the vocabulary garden — only words the dictionary
// actually knows (a real gloss or definition), so accidental lookups of
@@ -723,14 +953,18 @@ export function EditorCore({
// Reflect the saved state optimistically so the heart shows 💚 the
// moment a known word loads, rather than flashing 🤍 until the capture
// round-trips. vocabId is filled in when recordVocab returns.
setWordInfo((w) => (w ? { ...w, loading: false, info, saved: known } : null))
setWordInfo((w) => (w ? { ...w, loading: false, info, pinyin, saved: known } : null))
if (!known) return
api
.recordVocab({
word: range.word,
gloss: info.gloss,
definition: info.definitions[0]?.definition ?? '',
phonetic: info.phonetic,
// The garden's pronunciation field holds whichever this word has:
// IPA for an English word, pinyin for a Chinese one. Both answer
// the same question on a review card — how do I say this — and a
// second column would only be a second thing to keep in sync.
phonetic: pinyin || info.phonetic,
example,
doc_id: docId,
})
@@ -750,7 +984,7 @@ export function EditorCore({
}
})
},
[editor, closeCard, docId],
[editor, closeCard, docId, resolveWord, exampleAt],
)
// Toggle a looked-up word in/out of the vocabulary garden from the WordCard
@@ -797,13 +1031,16 @@ export function EditorCore({
if (!editor) return
const coords = editor.view.posAtCoords({ left: e.clientX, top: e.clientY })
if (!coords) return
if (!wordAt(editor.state.doc, coords.pos, wordAlphabet)) return
// Whichever script is under the pointer — the same resolver openWordLookup
// uses, so a Chinese word gets the card here too rather than falling
// through to the native menu.
if (!resolveWord(coords.pos)) return
e.preventDefault()
// A misspelled word offers corrections first; otherwise look it up.
if (openMisspellAt(coords.pos)) return
openWordLookup(coords.pos)
},
[editor, wordAlphabet, openMisspellAt, openWordLookup],
[editor, resolveWord, openMisspellAt, openWordLookup],
)
// Touch has no hover or right-click, so a long-press (~500ms without moving)
@@ -857,7 +1094,7 @@ export function EditorCore({
clear()
return
}
const range = wordAt(editor.state.doc, coords.pos, wordAlphabet)
const range = resolveWord(coords.pos)
if (!range) {
clear()
return
@@ -866,9 +1103,21 @@ export function EditorCore({
if (gloss && gloss.from === range.from && gloss.to === range.to) return
clearTimeout(glossTimer.current)
const token = ++glossReqRef.current
// The Chinese hover carries a second line the English one has no use for:
// pinyin above the meaning. It is the thing a learner most often stops to
// ask about their own writing — reading a character back is not the same
// as being able to say it — and it is why this tooltip is worth having at
// all for a script the writer can already read the meaning of half the
// time.
const ask = (): Promise<{ gloss: string; reverse?: string; lead?: string }> =>
range.hanzi
? api.hanziWord(range.word).then((h) => ({
gloss: h.readings[0]?.senses ?? h.chars.map((c) => `${c.char} ${c.senses}`).join(' · '),
lead: hanziPinyin(h),
}))
: api.glossWord(range.word).then((g) => ({ gloss: g.gloss, reverse: g.reverse }))
glossTimer.current = setTimeout(() => {
api
.glossWord(range.word)
ask()
.then((g) => {
if (token !== glossReqRef.current) return
const wrapper = wrapperRef.current
@@ -883,14 +1132,14 @@ export function EditorCore({
const wrapRect = wrapper.getBoundingClientRect()
const left = Math.max(0, Math.min(start.left - wrapRect.left, wrapper.clientWidth - 280))
const top = end.bottom - wrapRect.top + 6
setGloss({ word: range.word, gloss: g.gloss, reverse: g.reverse, from: range.from, to: range.to, top, left })
setGloss({ word: range.word, gloss: g.gloss, reverse: g.reverse, lead: g.lead, from: range.from, to: range.to, top, left })
})
.catch(() => {
if (token === glossReqRef.current) setGloss(null)
})
}, 350)
},
[editor, wordAlphabet, selection, rewrite, misspell, wordInfo, pinned, gloss],
[editor, resolveWord, selection, rewrite, misspell, wordInfo, pinned, gloss],
)
// Leaving the editor surface drops any pending/shown gloss.
@@ -1082,6 +1331,11 @@ export function EditorCore({
<div
ref={wrapperRef}
className="relative flex-1"
// Grown to cover whichever absolutely-positioned surface reaches lowest —
// the rail's card stack, or an open anchored card — so the space those
// cards occupy is actually scrollable. `minHeight` never shrinks the
// column, so a rail or card that fits beside its text changes nothing.
style={overhang > 0 ? { minHeight: overhang } : undefined}
onMouseOver={handleMouseOver}
onMouseOut={handleMouseOut}
onMouseMove={handleMouseMove}
@@ -1092,13 +1346,24 @@ export function EditorCore({
onTouchStart={handleTouchStart}
onTouchMove={cancelLongPress}
onTouchEnd={cancelLongPress}
>
{/* The prose sits in its own box so it can be measured (and pinned)
independently of the wrapper, which the rail may have grown. The box is
deliberately left at its natural height: sized to the wrapper it would
report the stack's height back as the text's own, and the pin below
could never trip. */}
<div
ref={contentRef}
style={stickTop === null ? undefined : { position: 'sticky', top: stickTop }}
>
<EditorContent editor={editor} className="h-full" />
</div>
{findOpen && editor && <FindReplace editor={editor} onClose={() => setFindOpen(false)} />}
{confetti && <Confetti top={confetti.top} left={confetti.left} />}
{gloss && (
<GlossTip
gloss={gloss.gloss}
lead={gloss.lead}
reverse={gloss.reverse}
reverseLang={pack.nativeName}
style={{ top: gloss.top, left: gloss.left }}
@@ -1131,6 +1396,7 @@ export function EditorCore({
loading={wordInfo.loading}
saved={wordInfo.saved}
onToggleSave={toggleSaveWord}
pinyin={wordInfo.pinyin}
style={{ top: wordInfo.top, left: wordInfo.left }}
onReplace={replaceWord}
/>
@@ -1144,15 +1410,18 @@ export function EditorCore({
onAdd={addMisspellingToDict}
/>
)}
{hover && !railEnabled && (
{hover && (
<SuggestionCard
suggestion={hover.suggestion}
style={{ top: hover.top, left: hover.left }}
batchCount={batchCounts[hover.suggestion.type] ?? 0}
onAccept={handleAccept}
onAcceptAll={handleAcceptAll}
onDismiss={handleDismiss}
onPointerEnter={keepOpen}
onPointerLeave={scheduleClose}
onExpandChange={setPinned}
onExtent={setCardExtent}
/>
)}
{railEnabled && railItems.length > 0 && (
@@ -1160,11 +1429,14 @@ export function EditorCore({
items={railItems}
activeId={activeId}
expandedId={railExpandedId}
batchCounts={batchCounts}
onAccept={handleAccept}
onAcceptAll={handleAcceptAll}
onDismiss={handleDismiss}
onHover={setActiveId}
onActivate={activateRailCard}
onToggleExpand={toggleRailExpand}
onExtent={setRailExtent}
/>
)}
</div>
+7 -3
View File
@@ -2,6 +2,7 @@ import { useCallback, useEffect, useRef, useState } from 'react'
import type { Editor } from '@tiptap/react'
import { clearSearch, getSearchState, setActive, setSearch } from './SearchHighlight'
import { usePack } from '../../i18n'
import { fromIME } from '../../lib/ime'
// FindReplace is the in-document search bar (Ctrl/Cmd+F). It drives the
// SearchHighlight decoration layer: typing updates the highlighted matches, the
@@ -108,7 +109,10 @@ export function FindReplace({ editor, onClose }: Props) {
role="dialog"
aria-label="Find and replace"
onKeyDown={(e) => {
if (e.key === 'Escape') {
// Both fields take Chinese, so both take an IME: Escape cancels a
// candidate and Enter commits one. A key that belongs to the composition
// is not a command here — see lib/ime.
if (e.key === 'Escape' && !fromIME(e)) {
e.preventDefault()
onClose()
}
@@ -137,7 +141,7 @@ export function FindReplace({ editor, onClose }: Props) {
value={query}
onChange={(e) => setQuery(e.target.value)}
onKeyDown={(e) => {
if (e.key === 'Enter') {
if (e.key === 'Enter' && !fromIME(e)) {
e.preventDefault()
go(e.shiftKey ? -1 : 1)
}
@@ -171,7 +175,7 @@ export function FindReplace({ editor, onClose }: Props) {
value={replacement}
onChange={(e) => setReplacement(e.target.value)}
onKeyDown={(e) => {
if (e.key === 'Enter') {
if (e.key === 'Enter' && !fromIME(e)) {
e.preventDefault()
replaceActive()
}
+11 -1
View File
@@ -7,6 +7,11 @@
interface Props {
gloss: string
// A line above the gloss, in a lighter weight: the pinyin of a Chinese word.
// It leads because it is what is actually being asked — a learner reading
// their own 公园 back may know it means a park and still not know how to say
// it, which is the one thing the character does not tell them.
lead?: string
// The English meaning of the same token read as a word of the writer's own
// language, when it is one. On a Latin-script pair "sale" is both, and the
// bubble shows the two readings stacked rather than picking one — the same
@@ -17,7 +22,7 @@ interface Props {
style: React.CSSProperties
}
export function GlossTip({ gloss, reverse, reverseLang, style }: Props) {
export function GlossTip({ gloss, lead, reverse, reverseLang, style }: Props) {
return (
<div
className="petal-gloss-tip pointer-events-none absolute z-20 px-2.5 py-1.5 text-sm"
@@ -33,6 +38,11 @@ export function GlossTip({ gloss, reverse, reverseLang, style }: Props) {
...style,
}}
>
{lead && (
<span className="mb-0.5 block font-semibold" style={{ opacity: 0.9 }}>
{lead}
</span>
)}
{gloss}
{reverse && (
<span className="mt-0.5 block" style={{ opacity: 0.72, fontSize: '0.85em' }}>
+45 -19
View File
@@ -4,6 +4,7 @@ import type { EditorState, Transaction } from '@tiptap/pm/state'
import { Decoration, DecorationSet } from '@tiptap/pm/view'
import type { Node as PMNode } from '@tiptap/pm/model'
import { mapOffset } from './SuggestionHighlight'
import { holdRedraw } from './Composition'
// SearchHighlight powers the in-document Find & Replace bar. Like the suggestion
// layer it uses ProseMirror *decorations* (not stored marks), so matches are
@@ -22,6 +23,11 @@ interface PluginState {
matches: Match[]
active: number // index into matches, or -1 when there are none
decorations: DecorationSet
// Held back while an IME composition was in flight — see Composition.ts.
// `matches` is held with the decorations rather than recomputed on its own:
// the Find bar's "3 / 7" and the wash on the page are one answer, and half of
// it moving while the other half waits would be worse than both waiting.
stale: boolean
}
export const searchPluginKey = new PluginKey<PluginState>('petalSearch')
@@ -57,7 +63,7 @@ function build(doc: PMNode, query: string, caseSensitive: boolean, preferred: nu
class: i === active ? 'petal-find-match petal-find-match-active' : 'petal-find-match',
}),
)
return { query, caseSensitive, matches, active, decorations: DecorationSet.create(doc, decos) }
return { query, caseSensitive, matches, active, decorations: DecorationSet.create(doc, decos), stale: false }
}
const EMPTY: PluginState = {
@@ -66,6 +72,7 @@ const EMPTY: PluginState = {
matches: [],
active: -1,
decorations: DecorationSet.empty,
stale: false,
}
// setSearch updates the query / case-sensitivity and recomputes matches. Passing
@@ -100,36 +107,49 @@ type Meta =
| { kind: 'active'; index: number }
| { kind: 'clear' }
export const SearchHighlight = Extension.create({
name: 'searchHighlight',
addProseMirrorPlugins() {
return [
new Plugin<PluginState>({
export function searchPlugin(): Plugin<PluginState> {
return new Plugin<PluginState>({
key: searchPluginKey,
state: {
init: () => EMPTY,
apply(tr, value, _oldState, newState): PluginState {
apply(tr, value, oldState, newState): PluginState {
const meta = tr.getMeta(searchPluginKey) as Meta | undefined
if (meta?.kind === 'search') {
return build(newState.doc, meta.query, meta.caseSensitive, value.active < 0 ? 0 : value.active)
}
// Clearing the layer is the one thing a composition never holds: it
// removes decorations rather than adding them, and it is what closing
// the Find bar does.
if (meta?.kind === 'clear') return EMPTY
const held = holdRedraw(tr, oldState)
const query = meta?.kind === 'search' ? meta.query : value.query
const caseSensitive = meta?.kind === 'search' ? meta.caseSensitive : value.caseSensitive
if (meta?.kind === 'active') {
if (value.matches.length === 0) return value
const active = ((meta.index % value.matches.length) + value.matches.length) % value.matches.length
if (held) return { ...value, active, stale: true }
const decos = value.matches.map((m, i) =>
Decoration.inline(m.from, m.to, {
class: i === active ? 'petal-find-match petal-find-match-active' : 'petal-find-match',
}),
)
return { ...value, active, decorations: DecorationSet.create(newState.doc, decos) }
return { ...value, active, stale: false, decorations: DecorationSet.create(newState.doc, decos) }
}
if (meta?.kind === 'clear') return EMPTY
// Re-anchor on any document change so highlights track edits/replaces.
if (tr.docChanged && value.query) {
return build(newState.doc, value.query, value.caseSensitive, value.active)
// A new query, or any document change: re-anchor so highlights track
// edits and replaces. Once due, it stays due until it happens.
const due = value.stale || meta?.kind === 'search' || (tr.docChanged && !!value.query)
if (!due) return value
if (held) {
return {
...value,
query,
caseSensitive,
stale: true,
decorations: tr.docChanged ? value.decorations.map(tr.mapping, tr.doc) : value.decorations,
}
return value
}
const preferred = meta?.kind === 'search' && value.active < 0 ? 0 : value.active
return build(newState.doc, query, caseSensitive, preferred)
},
},
props: {
@@ -137,7 +157,13 @@ export const SearchHighlight = Extension.create({
return searchPluginKey.getState(state)?.decorations
},
},
}),
]
})
}
export const SearchHighlight = Extension.create({
name: 'searchHighlight',
addProseMirrorPlugins() {
return [searchPlugin()]
},
})
+32 -14
View File
@@ -4,6 +4,7 @@ import type { EditorState, Transaction } from '@tiptap/pm/state'
import { Decoration, DecorationSet } from '@tiptap/pm/view'
import type { Node as PMNode } from '@tiptap/pm/model'
import { mapOffset } from './SuggestionHighlight'
import { holdRedraw } from './Composition'
import type { SpellChecker } from '../../hooks/useSpellChecker'
// SpellCheck renders browser-side nspell misspellings as ProseMirror
@@ -18,6 +19,11 @@ export const spellPluginKey = new PluginKey<PluginState>('petalSpellCheck')
interface PluginState {
checker: SpellChecker | null
decorations: DecorationSet
// Held back while an IME composition was in flight — see Composition.ts. This
// layer is the one with the most to gain from the guard: the pinyin she is
// part-way through typing is Latin letters, so it is exactly what the
// tokenizer picks up and exactly what an underline would redraw over.
stale: boolean
}
// A word is a run of Latin letters with optional internal/edge apostrophes
@@ -151,25 +157,31 @@ export function setSpellChecker(
dispatch(state.tr.setMeta(spellPluginKey, checker ?? null))
}
export const SpellCheck = Extension.create({
name: 'spellCheck',
addProseMirrorPlugins() {
return [
new Plugin<PluginState>({
export function spellPlugin(): Plugin<PluginState> {
return new Plugin<PluginState>({
key: spellPluginKey,
state: {
init: () => ({ checker: null, decorations: DecorationSet.empty }),
apply(tr, value, _oldState, newState) {
init: () => ({ checker: null, decorations: DecorationSet.empty, stale: false }),
apply(tr, value, oldState, newState) {
const meta = tr.getMeta(spellPluginKey) as SpellChecker | null | undefined
const checker = meta !== undefined ? meta : value.checker
if (!checker) return { checker: null, decorations: DecorationSet.empty }
if (!checker) return { checker: null, decorations: DecorationSet.empty, stale: false }
// Rebuild on a checker swap, a doc edit, or a caret move (so the word
// you just left gets re-evaluated and the new caret word is exempt).
if (meta !== undefined || tr.docChanged || tr.selectionSet) {
return { checker, decorations: buildDecorations(newState.doc, checker, newState.selection.head) }
const due = value.stale || meta !== undefined || tr.docChanged || tr.selectionSet
if (!due) return value
if (holdRedraw(tr, oldState)) {
return {
checker,
stale: true,
decorations: tr.docChanged ? value.decorations.map(tr.mapping, tr.doc) : value.decorations,
}
}
return {
checker,
stale: false,
decorations: buildDecorations(newState.doc, checker, newState.selection.head),
}
return { checker, decorations: value.decorations }
},
},
props: {
@@ -177,7 +189,13 @@ export const SpellCheck = Extension.create({
return spellPluginKey.getState(state)?.decorations
},
},
}),
]
})
}
export const SpellCheck = Extension.create({
name: 'spellCheck',
addProseMirrorPlugins() {
return [spellPlugin()]
},
})
+52 -5
View File
@@ -1,18 +1,30 @@
import { useState } from 'react'
import type { Suggestion } from '../../api/client'
import { useEffect, useRef, useState } from 'react'
import type { Suggestion, SuggestionType } from '../../api/client'
import { usePack } from '../../i18n'
import { AskPetal } from './AskPetal'
import { TYPE_META } from './suggestionMeta'
import { TYPE_META, batchLabel, typeLabel } from './suggestionMeta'
interface Props {
suggestion: Suggestion
style: React.CSSProperties
// How many pending suggestions share this card's type (including this one).
// Two or more offers to take the whole category in one step.
batchCount: number
onAccept: (s: Suggestion) => void
onAcceptAll: (type: SuggestionType) => void
onDismiss: (s: Suggestion) => void
onPointerEnter: () => void
onPointerLeave: () => void
// Pins the card open while the Ask Petal panel is expanded, so the chat isn't
// dismissed by the hover-close timer when the pointer drifts away.
onExpandChange: (expanded: boolean) => void
// How far the card reaches below the wrapper's top, in wrapper coordinates —
// the same report the rail makes (item 4). The card is absolutely positioned
// and so adds no layout height of its own; without this, an Ask Petal
// conversation that runs past the last line of text has no scrollable page
// under it and its Accept button simply can't be reached. 0 means "nothing to
// cover", which is what an unmounted card reports on its way out.
onExtent?: (bottom: number) => void
}
// SuggestionCard is the hover panel for a single suggestion: a colored type tag,
@@ -22,15 +34,38 @@ interface Props {
export function SuggestionCard({
suggestion,
style,
batchCount,
onAccept,
onAcceptAll,
onDismiss,
onPointerEnter,
onPointerLeave,
onExpandChange,
onExtent,
}: Props) {
const pack = usePack()
const meta = TYPE_META[suggestion.type]
const label = typeLabel(suggestion.type, pack)
const hasReplacement = suggestion.replacement.trim() !== ''
const [asking, setAsking] = useState(false)
const cardRef = useRef<HTMLDivElement>(null)
// Report the card's reach while it is open, and withdraw it on the way out.
// A ResizeObserver rather than a one-shot measure because the card grows
// after it is mounted: the Ask Petal panel opens, and then the reply streams
// into it token by token.
useEffect(() => {
const el = cardRef.current
if (!el || !onExtent) return
const report = () => onExtent(el.offsetTop + el.offsetHeight)
report()
const observer = new ResizeObserver(report)
observer.observe(el)
return () => {
observer.disconnect()
onExtent(0)
}
}, [onExtent])
function toggleAsking() {
setAsking((prev) => {
@@ -42,8 +77,9 @@ export function SuggestionCard({
return (
<div
ref={cardRef}
role="dialog"
aria-label={`${meta.label} suggestion`}
aria-label={`${label} suggestion`}
onMouseEnter={onPointerEnter}
onMouseLeave={onPointerLeave}
className="petal-suggestion-card absolute z-20 p-3.5 text-sm"
@@ -60,7 +96,7 @@ export function SuggestionCard({
className="inline-flex items-center gap-1.5 rounded-full px-2.5 py-0.5 text-xs font-bold"
style={{ background: meta.color, color: 'var(--color-plum)' }}
>
{meta.label}
{label}
</span>
{hasReplacement && (
@@ -114,6 +150,17 @@ export function SuggestionCard({
Dismiss
</button>
</div>
{hasReplacement && batchCount > 1 && (
<button
type="button"
onClick={() => onAcceptAll(suggestion.type)}
className="petal-accept-all mt-2 w-full rounded-full py-1.5 text-xs font-bold"
style={{ color: 'var(--color-plum)', borderColor: meta.color }}
>
{batchLabel(suggestion.type, batchCount)}
</button>
)}
</div>
)
}
@@ -4,6 +4,7 @@ import type { EditorState, Transaction } from '@tiptap/pm/state'
import { Decoration, DecorationSet } from '@tiptap/pm/view'
import type { Node as PMNode } from '@tiptap/pm/model'
import type { Suggestion } from '../../api/client'
import { holdRedraw } from './Composition'
// SuggestionHighlight renders LLM suggestions as ProseMirror *decorations*, not
// stored marks. Decorations are ephemeral overlays recomputed from the live
@@ -21,6 +22,10 @@ interface PluginState {
// decoration repaints that fire on every document change.
activeId: string | null
decorations: DecorationSet
// A rebuild fell due while an IME composition was in flight and was held back
// (see Composition.ts). The decorations on screen are the previous ones,
// mapped forward; this says they still owe a rebuild.
stale: boolean
}
// Meta carried on a transaction to update the plugin: either a fresh suggestion
@@ -173,40 +178,37 @@ export function setActiveSuggestion(
dispatch(state.tr.setMeta(suggestionPluginKey, { activeId } satisfies SuggestionMeta))
}
export const SuggestionHighlight = Extension.create({
name: 'suggestionHighlight',
addProseMirrorPlugins() {
return [
new Plugin<PluginState>({
export function suggestionPlugin(): Plugin<PluginState> {
return new Plugin<PluginState>({
key: suggestionPluginKey,
state: {
init: () => ({ suggestions: [], activeId: null, decorations: DecorationSet.empty }),
apply(tr, value, _oldState, newState) {
init: () => ({ suggestions: [], activeId: null, decorations: DecorationSet.empty, stale: false }),
apply(tr, value, oldState, newState) {
const meta = tr.getMeta(suggestionPluginKey) as SuggestionMeta | undefined
if (meta && 'suggestions' in meta) {
const suggestions = meta && 'suggestions' in meta ? meta.suggestions : value.suggestions
const activeId = meta && 'activeId' in meta ? meta.activeId : value.activeId
// A rebuild is due on a new list, a new emphasis, or any document change
// (which is how a suggestion re-anchors by string), and stays due until
// it happens.
const due = value.stale || meta !== undefined || tr.docChanged
if (!due) return value
if (holdRedraw(tr, oldState)) {
return {
suggestions: meta.suggestions,
activeId: value.activeId,
decorations: buildDecorations(newState.doc, meta.suggestions, value.activeId),
suggestions,
activeId,
stale: true,
// Map rather than keep: the composing text is growing under these
// highlights, and an unmapped decoration would drift a character at
// a time across a word she is still typing.
decorations: tr.docChanged ? value.decorations.map(tr.mapping, tr.doc) : value.decorations,
}
}
if (meta && 'activeId' in meta) {
return {
suggestions: value.suggestions,
activeId: meta.activeId,
decorations: buildDecorations(newState.doc, value.suggestions, meta.activeId),
suggestions,
activeId,
stale: false,
decorations: buildDecorations(newState.doc, suggestions, activeId),
}
}
// On any document change, re-anchor by string against the new doc.
if (tr.docChanged) {
return {
suggestions: value.suggestions,
activeId: value.activeId,
decorations: buildDecorations(newState.doc, value.suggestions, value.activeId),
}
}
return value
},
},
props: {
@@ -214,7 +216,13 @@ export const SuggestionHighlight = Extension.create({
return suggestionPluginKey.getState(state)?.decorations
},
},
}),
]
})
}
export const SuggestionHighlight = Extension.create({
name: 'suggestionHighlight',
addProseMirrorPlugins() {
return [suggestionPlugin()]
},
})
+56 -7
View File
@@ -1,7 +1,9 @@
import { forwardRef, useLayoutEffect, useRef, useState } from 'react'
import type { Suggestion } from '../../api/client'
import type { Suggestion, SuggestionType } from '../../api/client'
import { usePack } from '../../i18n'
import { AskPetal } from './AskPetal'
import { TYPE_META } from './suggestionMeta'
import { batchLeaders } from './acceptBatch'
import { TYPE_META, batchLabel, typeLabel } from './suggestionMeta'
// Vertical breathing room kept between stacked cards when their natural anchors
// would otherwise collide.
@@ -21,13 +23,23 @@ interface Props {
activeId: string | null
// The card expanded to show the full explanation + Ask Petal, or null.
expandedId: string | null
// Pending suggestions per type. A card whose type has company offers to take
// the whole category in one step (and one undo step).
batchCounts: Partial<Record<SuggestionType, number>>
onAccept: (s: Suggestion) => void
onAcceptAll: (type: SuggestionType) => void
onDismiss: (s: Suggestion) => void
// Pointer entering/leaving a card, so the matching highlight can light up.
onHover: (id: string | null) => void
// A card's body was clicked — scroll its highlight into view and toggle expand.
onActivate: (id: string) => void
onToggleExpand: (id: string) => void
// How far down the resolved stack reaches (px below the wrapper's top). Cards
// are absolutely positioned, so they add nothing to layout height — a cluster of
// errors in one short paragraph can pile cards hundreds of px past the end of the
// text, with no scrollable space to reach them. The editor uses this to grow the
// column so every card can at least be scrolled to.
onExtent: (bottom: number) => void
}
// SuggestionRail is the right-margin "comment column": every outstanding
@@ -39,11 +51,14 @@ export function SuggestionRail({
items,
activeId,
expandedId,
batchCounts,
onAccept,
onAcceptAll,
onDismiss,
onHover,
onActivate,
onToggleExpand,
onExtent,
}: Props) {
// Measured resolved tops keyed by suggestion id (after collision avoidance).
const [tops, setTops] = useState<Record<string, number>>({})
@@ -65,13 +80,16 @@ export function SuggestionRail({
const layoutKey = ordered.map((i) => `${i.suggestion.id}:${Math.round(i.anchorTop)}`).join('|')
useLayoutEffect(() => {
let cursor = -Infinity
let bottom = 0
const next: Record<string, number> = {}
for (const { suggestion, anchorTop } of ordered) {
const h = cardRefs.current.get(suggestion.id)?.offsetHeight ?? 96
const top = Math.max(anchorTop, cursor)
next[suggestion.id] = top
cursor = top + h + CARD_GAP
bottom = top + h
}
onExtent(bottom)
setTops((prev) => {
const ids = Object.keys(next)
if (ids.length === Object.keys(prev).length && ids.every((id) => prev[id] === next[id])) return prev
@@ -82,6 +100,9 @@ export function SuggestionRail({
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [layoutKey, expandedId, measureTick])
// One card per category carries the batch control — see batchLeaders.
const batchLead = batchLeaders(ordered.map(({ suggestion }) => suggestion))
return (
<div className="petal-rail petal-no-print" aria-label="Suggestions">
{ordered.map(({ suggestion }) => (
@@ -101,7 +122,9 @@ export function SuggestionRail({
top={tops[suggestion.id] ?? 0}
active={activeId === suggestion.id}
expanded={expandedId === suggestion.id}
batchCount={batchLead.has(suggestion.id) ? (batchCounts[suggestion.type] ?? 0) : 0}
onAccept={onAccept}
onAcceptAll={onAcceptAll}
onDismiss={onDismiss}
onHover={onHover}
onActivate={onActivate}
@@ -117,7 +140,9 @@ interface CardProps {
top: number
active: boolean
expanded: boolean
batchCount: number
onAccept: (s: Suggestion) => void
onAcceptAll: (type: SuggestionType) => void
onDismiss: (s: Suggestion) => void
onHover: (id: string | null) => void
onActivate: (id: string) => void
@@ -125,17 +150,24 @@ interface CardProps {
}
const RailCard = forwardRef<HTMLDivElement, CardProps>(function RailCard(
{ suggestion, top, active, expanded, onAccept, onDismiss, onHover, onActivate, onToggleExpand },
{ suggestion, top, active, expanded, batchCount, onAccept, onAcceptAll, onDismiss, onHover, onActivate, onToggleExpand },
ref,
) {
const pack = usePack()
const meta = TYPE_META[suggestion.type]
const label = typeLabel(suggestion.type, pack)
const hasReplacement = suggestion.replacement.trim() !== ''
// Every other card truncates its diff to one line each: the original and the
// replacement differ by a word or two, and the explanation below is the part
// she reads. A translation is the reverse — the two lines are a whole sentence
// in each language, and they are the entire point of the card — so it wraps.
const clampDiff = suggestion.type !== 'translate'
return (
<div
ref={ref}
role="group"
aria-label={`${meta.label} suggestion`}
aria-label={`${label} suggestion`}
onMouseEnter={() => onHover(suggestion.id)}
onMouseLeave={() => onHover(null)}
className={`petal-rail-card${active ? ' petal-rail-card-active' : ''}`}
@@ -146,7 +178,7 @@ const RailCard = forwardRef<HTMLDivElement, CardProps>(function RailCard(
className="inline-flex items-center rounded-full px-2 py-0.5 text-[0.68rem] font-bold"
style={{ background: meta.color, color: 'var(--color-plum)' }}
>
{meta.label}
{label}
</span>
<button
type="button"
@@ -163,10 +195,16 @@ const RailCard = forwardRef<HTMLDivElement, CardProps>(function RailCard(
<button type="button" onClick={() => onActivate(suggestion.id)} className="mt-2 block w-full text-left">
{hasReplacement && (
<span className="flex flex-col gap-0.5" style={{ fontFamily: 'var(--font-body)' }}>
<span className="truncate text-[0.9rem] line-through" style={{ color: 'var(--color-muted)' }}>
<span
className={`text-[0.9rem] line-through${clampDiff ? ' truncate' : ''}`}
style={{ color: 'var(--color-muted)' }}
>
{suggestion.original}
</span>
<span className="truncate text-[0.9rem] font-medium" style={{ color: 'var(--color-plum)' }}>
<span
className={`text-[0.9rem] font-medium${clampDiff ? ' truncate' : ''}`}
style={{ color: 'var(--color-plum)' }}
>
{suggestion.replacement}
</span>
</span>
@@ -206,6 +244,17 @@ const RailCard = forwardRef<HTMLDivElement, CardProps>(function RailCard(
{expanded ? 'Hide Petal' : 'Ask Petal ✨'}
</button>
</div>
{hasReplacement && batchCount > 1 && (
<button
type="button"
onClick={() => onAcceptAll(suggestion.type)}
className="petal-accept-all mt-2 w-full rounded-full py-1 text-[0.7rem] font-bold"
style={{ color: 'var(--color-plum)', borderColor: meta.color }}
>
{batchLabel(suggestion.type, batchCount)}
</button>
)}
</div>
)
})
+9 -3
View File
@@ -17,11 +17,16 @@ interface Props {
// heart toggles it; `onToggleSave` removes/re-adds it.
saved: boolean
onToggleSave: () => void
// A Chinese word's pinyin. Shown in place of the IPA line and *without* the
// slashes, because pinyin is not a phonetic transcription — it is how the word
// is spelled in letters, and the slashes would say something untrue about it
// in the one place a learner is looking for the truth about pronunciation.
pinyin?: string
style: React.CSSProperties
onReplace: (synonym: string) => void
}
export function WordCard({ word, info, loading, saved, onToggleSave, style, onReplace }: Props) {
export function WordCard({ word, info, loading, saved, onToggleSave, pinyin, style, onReplace }: Props) {
const t = usePack()
const definitions = info?.definitions ?? []
const synonyms = info?.synonyms ?? []
@@ -117,9 +122,10 @@ export function WordCard({ word, info, loading, saved, onToggleSave, style, onRe
when she has found a word she likes "can I use this?". Both are
quiet, muted lines: information she can take or leave, never a verdict
on her writing. */}
{(phonetic || band) && (
{(phonetic || pinyin || band) && (
<div className="mt-1.5 flex items-center gap-2 text-sm">
{phonetic && <span style={{ color: 'var(--color-muted)' }}>/{phonetic}/</span>}
{pinyin && <span style={{ color: 'var(--color-muted)' }}>{pinyin}</span>}
{!pinyin && phonetic && <span style={{ color: 'var(--color-muted)' }}>/{phonetic}/</span>}
{band && (
<span
className="rounded-full px-2 py-0.5 text-xs font-semibold"
@@ -0,0 +1,128 @@
import { describe, expect, it } from 'vitest'
import type { Suggestion, SuggestionType } from '../../api/client'
import { batchLeaders, planBatch } from './acceptBatch'
function sug(id: string, original: string, replacement: string, type: SuggestionType = 'grammar'): Suggestion {
return {
id,
doc_id: 'd1',
from_pos: 0,
to_pos: 0,
original,
replacement,
explanation: '',
type,
status: 'pending',
created_at: '2026-07-28T00:00:00Z',
}
}
// A stand-in for findRange: first occurrence in a flat string, or null.
function finder(text: string) {
return (needle: string) => {
const i = text.indexOf(needle)
return i < 0 ? null : { from: i, to: i + needle.length }
}
}
describe('planBatch', () => {
it('applies from the end backwards, so earlier spans keep their positions', () => {
const text = 'a apple and a orange and a egg'
const plan = planBatch(
[sug('1', 'a apple', 'an apple'), sug('2', 'a orange', 'an orange'), sug('3', 'a egg', 'an egg')],
finder(text),
)
expect(plan.steps.map((s) => s.suggestion.id)).toEqual(['3', '2', '1'])
expect(plan.missing).toEqual([])
expect(plan.skipped).toEqual([])
// Applying the plan in order against a mutable string must land every
// replacement where it belongs — this is the property the ordering exists
// for, and the one a single shared transaction depends on.
let out = text
for (const step of plan.steps) {
out = out.slice(0, step.from) + step.suggestion.replacement + out.slice(step.to)
}
expect(out).toBe('an apple and an orange and an egg')
})
it('orders by position in the document, not by the order the cards arrived', () => {
const plan = planBatch(
[sug('late', 'a egg', 'an egg'), sug('early', 'a apple', 'an apple')],
finder('a apple and a egg'),
)
expect(plan.steps.map((s) => s.suggestion.id)).toEqual(['late', 'early'])
})
it('settles a card whose span is gone — she already fixed it herself', () => {
const plan = planBatch(
[sug('1', 'a apple', 'an apple'), sug('2', 'a orange', 'an orange')],
finder('an apple and a orange'),
)
expect(plan.steps.map((s) => s.suggestion.id)).toEqual(['2'])
expect(plan.missing.map((s) => s.id)).toEqual(['1'])
})
it('leaves the second of two cards quoting the same words on screen', () => {
// findRange resolves both to the first occurrence, so applying both would
// overwrite the first edit with the second. The batch takes one and says so.
const plan = planBatch(
[sug('1', 'a apple', 'an apple'), sug('2', 'a apple', 'the apple')],
finder('a apple, a apple'),
)
expect(plan.steps.map((s) => s.suggestion.id)).toEqual(['1'])
expect(plan.skipped.map((s) => s.id)).toEqual(['2'])
expect(plan.missing).toEqual([])
})
it('keeps both of two spans that merely touch', () => {
// Adjacent is not overlapping: "to" ends exactly where "the" begins.
const plan = planBatch(
[sug('1', 'aa', 'AA'), sug('2', 'bb', 'BB')],
finder('aabb'),
)
expect(plan.steps.map((s) => s.suggestion.id)).toEqual(['2', '1'])
expect(plan.skipped).toEqual([])
})
it('skips an awareness-only card with nothing to insert', () => {
const plan = planBatch([sug('v', 'a apple', ' ', 'voice')], finder('a apple'))
expect(plan.steps).toEqual([])
expect(plan.skipped.map((s) => s.id)).toEqual(['v'])
})
it('returns an empty plan for an empty family', () => {
const plan = planBatch([], finder('anything'))
expect(plan).toEqual({ steps: [], missing: [], skipped: [] })
})
})
describe('batchLeaders', () => {
const card = (id: string, type: SuggestionType) => ({ id, type })
it('gives each category exactly one leader, the first of its kind', () => {
const leaders = batchLeaders([
card('g1', 'grammar'),
card('m1', 'mechanics'),
card('g2', 'grammar'),
card('m2', 'mechanics'),
card('g3', 'grammar'),
])
expect([...leaders]).toEqual(['g1', 'm1'])
})
it('follows the stacking order it is given, not the order types appear elsewhere', () => {
// The rail sorts by anchor, so "first" means highest on the page — the card
// she reads first, which is where the batch control belongs.
const leaders = batchLeaders([card('lower', 'grammar'), card('upper', 'grammar')])
expect([...leaders]).toEqual(['lower'])
})
it('leads a lone card of its type too — the count is what hides the button', () => {
expect([...batchLeaders([card('only', 'voice')])]).toEqual(['only'])
})
it('has no leaders for an empty stack', () => {
expect(batchLeaders([]).size).toBe(0)
})
})
+84
View File
@@ -0,0 +1,84 @@
import type { Suggestion, SuggestionType } from '../../api/client'
// batchLeaders picks the one card per category that carries the "Accept all
// Grammar (5)" control, given the cards in the order they are stacked. The rail
// can't carry a category header — cards are anchored to their own sentence, so a
// type's cards are scattered down the column — and the closest honest stand-in
// is the first card of its kind. Offering the same batch on all five would be
// five buttons saying one thing.
export function batchLeaders(ordered: { type: SuggestionType; id: string }[]): Set<string> {
const leaders = new Set<string>()
const claimed = new Set<SuggestionType>()
for (const { type, id } of ordered) {
if (claimed.has(type)) continue
claimed.add(type)
leaders.add(id)
}
return leaders
}
export interface BatchStep {
suggestion: Suggestion
from: number
to: number
}
export interface BatchPlan {
// The edits to apply, ordered LAST span first. Every position is resolved
// against the document as it stands *before* any of them are applied, and
// applying from the end backwards means an earlier span's position can't be
// shifted by a later one — so the whole batch can go into a single
// transaction, which is the point: one Ctrl+Z puts it all back.
steps: BatchStep[]
// Cards whose span is no longer in the document — she fixed it herself, or
// edited around it since the check ran. There is nothing to replace, but the
// card is stale and settling it is what a single Accept already does.
missing: Suggestion[]
// Cards deliberately left on screen: two suggestions quoting the same words
// (findRange resolves both to the first occurrence, so applying the second
// would overwrite the first), and anything with nothing to insert. A batch
// that silently dropped these would report edits it never made.
skipped: Suggestion[]
}
// planBatch decides what one "Accept all <category>" click actually does. It is
// pure and takes the span resolver as an argument so the ordering and overlap
// rules can be tested without a ProseMirror document — the part that goes wrong
// is the arithmetic, not the lookup.
export function planBatch(
list: Suggestion[],
resolve: (original: string) => { from: number; to: number } | null,
): BatchPlan {
const located: BatchStep[] = []
const missing: Suggestion[] = []
const skipped: Suggestion[] = []
for (const suggestion of list) {
if (suggestion.replacement.trim() === '') {
skipped.push(suggestion) // awareness-only (voice): nothing to accept
continue
}
const range = resolve(suggestion.original)
if (!range) {
missing.push(suggestion)
continue
}
located.push({ suggestion, from: range.from, to: range.to })
}
// Document order, then greedily keep the non-overlapping ones.
located.sort((a, b) => a.from - b.from || a.to - b.to)
const steps: BatchStep[] = []
let cursor = -Infinity
for (const step of located) {
if (step.from < cursor) {
skipped.push(step.suggestion)
continue
}
steps.push(step)
cursor = step.to
}
steps.reverse()
return { steps, missing, skipped }
}
@@ -0,0 +1,74 @@
import { describe, expect, it } from 'vitest'
import { splitBilingual } from './bilingualReply'
// The contract these tests defend is "never hide an answer", not "parse the
// model". Every case that isn't cleanly two halves must still come back whole.
describe('splitBilingual', () => {
it('splits the pair language from the English at the blank line', () => {
const { native, en } = splitBilingual(
'“by foots” 不是固定说法,正确的是 “on foot”。\n\n"By foots" isnt a set phrase — the idiom is "on foot".',
)
expect(native).toBe('“by foots” 不是固定说法,正确的是 “on foot”。')
expect(en).toBe('"By foots" isnt a set phrase — the idiom is "on foot".')
})
it('works the same for a Latin pair, where both halves are Latin script', () => {
const { native, en } = splitBilingual(
'Dizemos "on foot", não "by foots".\n\nWe say "on foot", not "by foots".',
)
expect(native).toBe('Dizemos "on foot", não "by foots".')
expect(en).toBe('We say "on foot", not "by foots".')
})
it('renders a half-streamed reply as the pair language until the break arrives', () => {
// Mid-stream: the English half hasn't been written yet. The partial text is
// the whole bubble, not an empty one.
expect(splitBilingual('“by foots” 不是固定')).toEqual({
native: '“by foots” 不是固定',
en: '',
})
})
it('keeps a one-language reply whole', () => {
// A model that ignores the instruction costs styling, never content.
const single = 'We say "on foot" because the idiom is fixed.'
expect(splitBilingual(single)).toEqual({ native: single, en: '' })
})
it('treats extra blank lines as part of the English half', () => {
const { native, en } = splitBilingual('中文回答。\n\nFirst English point.\n\nSecond one.')
expect(native).toBe('中文回答。')
expect(en).toBe('First English point.\n\nSecond one.')
})
it('does not split on a blank line with nothing on one side', () => {
// A leading or trailing stray newline is not a separator; styling half of
// this as a translation of nothing would be worse than not splitting.
expect(splitBilingual('\n\nWe say "on foot".')).toEqual({
native: 'We say "on foot".',
en: '',
})
expect(splitBilingual('We say "on foot".\n\n')).toEqual({
native: 'We say "on foot".',
en: '',
})
})
it('accepts a separator line that carries whitespace', () => {
// Models emit "\n \n" often enough that requiring a bare "\n\n" would drop
// the split for a reply that followed the instruction.
const { native, en } = splitBilingual('中文回答。\n \nThe English answer.')
expect(native).toBe('中文回答。')
expect(en).toBe('The English answer.')
})
it('handles an empty reply', () => {
expect(splitBilingual('')).toEqual({ native: '', en: '' })
})
it('leaves single newlines inside a half alone', () => {
const { native, en } = splitBilingual('第一行\n第二行\n\nLine one\nLine two')
expect(native).toBe('第一行\n第二行')
expect(en).toBe('Line one\nLine two')
})
})
@@ -0,0 +1,43 @@
// Splitting Petal's chat reply into the two languages it was asked for.
//
// The Ask Petal prompt (internal/llm/prompts.go) asks for the pair language
// first, then the same answer in English, separated by one blank line. This is
// the reader of that contract — and it is deliberately forgiving, because the
// reply arrives from a small local model, token by token, and a rendering rule
// must never be able to hide an answer the writer could otherwise read.
//
// So there is exactly one failure mode and it is benign: anything that doesn't
// look like two halves is returned as `native` alone, which renders as one
// ordinary block. Nothing is dropped, ever.
export interface BilingualReply {
/** The pair language — or the whole reply, when there is only one half. */
native: string
/** The English half; '' when the reply hasn't reached the blank line yet. */
en: string
}
/**
* splitBilingual divides a reply at its first blank line.
*
* Streaming is the reason this splits at the *first* blank line rather than
* validating the shape: while tokens arrive the text is a native half with no
* separator yet, so it renders as the pair language and the English simply
* appears beneath it when the blank line lands. Any further blank lines stay
* inside the English half rather than starting a third section with nowhere to
* go.
*/
export function splitBilingual(content: string): BilingualReply {
const match = /\n[ \t]*\n/.exec(content)
if (!match) return { native: content, en: '' }
const native = content.slice(0, match.index).trim()
const en = content.slice(match.index + match[0].length).trim()
// A blank line with nothing on one side of it isn't two halves — it's a
// stray newline. Keep the reply whole rather than styling half of it as a
// translation of nothing.
if (native === '' || en === '') return { native: content.trim(), en: '' }
return { native, en }
}
+126
View File
@@ -0,0 +1,126 @@
import type { Node as PMNode } from '@tiptap/pm/model'
import { mapOffset } from './SuggestionHighlight'
import type { Segmenter } from '../../lib/segment'
import type { HanziInfo, WordInfo } from '../../api/client'
// The Chinese counterpart of `wordAt` (SpellCheck.ts): given a position in the
// document, which Chinese word is there.
//
// It lives in its own file rather than as a branch inside `wordAt` because the
// two answer the same question by genuinely different means — one runs a regex
// over the text, the other runs a shortest-path walk over a 188k-word list it
// had to fetch — and only one of them is about spelling at all. What they share
// is the part that matters for correctness: the offset→position mapping, which
// is `mapOffset`, the same function the suggestion, spell and search decoration
// layers all anchor through.
export interface HanziRange {
from: number
to: number
word: string
}
// blockAt finds the textblock containing pos, along with where that block starts
// — everything else here is arithmetic within one block.
//
// Segmentation is per-block for the same reason the spell tokenizer is: a word
// cannot span a paragraph break, and a block is the largest unit whose text is
// contiguous in the document.
function blockAt(doc: PMNode, pos: number): { node: PMNode; start: number } | null {
let found: { node: PMNode; start: number } | null = null
doc.descendants((node, nodePos) => {
if (found) return false
if (!node.isTextblock) return true
if (pos <= nodePos || pos >= nodePos + node.nodeSize) return false
found = { node, start: nodePos }
return false
})
return found
}
// offsetOf is the inverse of mapOffset: an absolute ProseMirror position to a
// character offset within the block's flattened text. Inline atoms (a hard
// break) occupy a position and contribute no text, so the two are not the same
// number and subtracting the block position would be wrong in any paragraph
// containing one.
function offsetOf(block: PMNode, blockStart: number, pos: number): number {
let textOffset = 0
let pmPos = blockStart + 1
let result = -1
block.forEach((child) => {
if (result >= 0) return
const len = child.isText ? (child.text?.length ?? 0) : 0
if (pos <= pmPos + child.nodeSize) {
result = textOffset + Math.max(0, Math.min(pos - pmPos, len))
return
}
textOffset += len
pmPos += child.nodeSize
})
return result >= 0 ? result : textOffset
}
// hanziWordAt resolves the Chinese word at a document position, or null when
// there is no Chinese there — which is the ordinary case in a mixed paragraph
// and is why the caller falls through to the Latin tokenizer.
export function hanziWordAt(doc: PMNode, pos: number, segmenter: Segmenter | null): HanziRange | null {
if (!segmenter) return null
const block = blockAt(doc, pos)
if (!block) return null
const text = block.node.textContent
if (!text) return null
const token = segmenter.wordAt(text, offsetOf(block.node, block.start, pos))
if (!token) return null
return {
from: mapOffset(block.node, block.start, token.from),
to: mapOffset(block.node, block.start, token.to),
word: token.word,
}
}
// hanziToWordInfo adapts a Chinese lookup into the shape the word card already
// renders.
//
// An adapter rather than a second card, because everything around the card is
// the same in both directions: it opens the same way, anchors the same way,
// captures into the same vocabulary garden, and reads aloud through the same
// voice — the zh pair already speaks Chinese, so 🔊 needs nothing new to say
// 公园 out loud. What differs is only which fields carry what.
//
// * `definitions` holds one entry per reading, labelled with its pinyin. The
// part-of-speech slot is where the card puts a short italic prefix, which
// is exactly the shape a reading label wants — and a reading *is* the thing
// that distinguishes these senses from each other (得 dé "to obtain" from 得
// de, the complement marker).
// * `gloss` stays empty. It means "translated into the writer's language",
// and for a writer learning Chinese that language is English, which is what
// the senses already are. Putting the English there too would print it
// twice.
// * The per-character fallback fills the same list, labelled by character, so
// a compound with no headword still says something true about itself.
export function hanziToWordInfo(info: HanziInfo): WordInfo {
const definitions =
info.readings.length > 0
? info.readings.map((r) => ({ part_of_speech: r.pinyin, definition: r.senses }))
: info.chars.map((c) => ({ part_of_speech: `${c.char} ${c.pinyin}`, definition: c.senses }))
return {
word: info.word,
gloss: '',
phonetic: '',
definitions,
synonyms: [],
frequency: 0,
difficulty: -1,
etymology: '',
}
}
// hanziPinyin is the word's own pronunciation, for the line under the headword.
// Empty when only the character fallback answered: the characters' readings are
// not the word's reading — 不 is bù alone and bú before a fourth tone — and
// printing them joined up would be inventing a pronunciation.
export function hanziPinyin(info: HanziInfo): string {
return info.readings[0]?.pinyin ?? ''
}
+42
View File
@@ -0,0 +1,42 @@
import { describe, expect, it } from 'vitest'
import { RAIL_MIN_MARGIN, railFitsBeside } from './railFit'
// The numbers below are not invented: they were measured in Chrome at the UX
// review's own 1517x810 viewport, with the 280px sidebar open and closed. They
// are here so the two layouts stay distinguishable if the constant is ever tuned.
describe('railFitsBeside', () => {
it('fits in distraction-free mode, where the pane spans the window', () => {
// 1517px window, sidebar collapsed: the 720px column centres at left 391,
// so its right edge is 1111 and 406px of margin remain.
expect(railFitsBeside(1517, 1111)).toBe(true)
})
it('does not fit with the document list open at the same window size', () => {
// Same window, 280px sidebar in flow: the column re-centres to right 1259 and
// the margin falls to 258 — the measurement item 5 reported. This is the case
// that must return false; rendering the rail here overhangs the viewport.
expect(railFitsBeside(1517, 1259)).toBe(false)
})
it('rejects the mid-animation width too, not just the settled one', () => {
// The sidebar animates over 280ms, so the recompute can land on an
// intermediate margin (266 was observed one frame in). Anything under the
// threshold has to read as "no rail", or the column flickers back in.
expect(railFitsBeside(1517, 1251)).toBe(false)
})
it('treats the threshold as inclusive', () => {
expect(railFitsBeside(1000, 1000 - RAIL_MIN_MARGIN)).toBe(true)
expect(railFitsBeside(1000, 1000 - RAIL_MIN_MARGIN + 1)).toBe(false)
})
it('has room to spare on a wide desktop', () => {
// 1920px maximised, sidebar open: margin 468.
expect(railFitsBeside(1920, 1452)).toBe(true)
})
it('never fits on a narrow window, whatever the column does', () => {
expect(railFitsBeside(900, 810)).toBe(false)
expect(railFitsBeside(768, 744)).toBe(false)
})
})
+19
View File
@@ -0,0 +1,19 @@
// Whether the margin rail has room to sit beside the editor.
//
// The rail is a 300px column with a 32px gutter (see `.petal-rail` in index.css);
// RAIL_MIN_MARGIN adds a little breathing space to the viewport edge. Below it the
// editor falls back to the inline card anchored under the word.
//
// This is a bare comparison, but it earns a name: the number decides which of two
// entirely different suggestion surfaces she gets, and the margin it measures moves
// for reasons that have nothing to do with the window size. The editor is a fixed
// 720px column centred in the pane, so collapsing the 280px sidebar (distraction-free
// mode) re-centres it and changes this margin by 140px without resizing anything.
// See the ResizeObserver in EditorCore for the other half of that story.
export const RAIL_MIN_MARGIN = 348
// `wrapperRight` and `innerWidth` are both viewport coordinates — i.e. exactly
// `wrapper.getBoundingClientRect().right` and `window.innerWidth`.
export function railFitsBeside(innerWidth: number, wrapperRight: number): boolean {
return innerWidth - wrapperRight >= RAIL_MIN_MARGIN
}
@@ -0,0 +1,54 @@
import { describe, expect, it } from 'vitest'
import type { SuggestionType } from '../../api/client'
import { fr } from '../../i18n/packs/fr'
import { ptPT } from '../../i18n/packs/pt-PT'
import { zh } from '../../i18n/packs/zh'
import { TYPE_META, typeLabel } from './suggestionMeta'
// Every type the API can send. Listed by hand rather than derived, so adding a
// suggestion type to the union without giving it a colour and a name fails here
// instead of rendering a card with `undefined` on its pill.
const ALL: SuggestionType[] = [
'grammar',
'phrasing',
'idiom',
'clarity',
'translate',
'voice',
'collocation',
'mechanics',
]
describe('TYPE_META', () => {
it('covers every suggestion type with a colour and a label', () => {
for (const type of ALL) {
expect(TYPE_META[type], type).toBeDefined()
expect(TYPE_META[type].color, `${type} colour`).toMatch(/^var\(--color-/)
expect(TYPE_META[type].label, `${type} label`).not.toBe('')
}
})
it('gives translate its own colour', () => {
const used = ALL.filter((t) => t !== 'translate').map((t) => TYPE_META[t].color)
expect(used).not.toContain(TYPE_META.translate.color)
})
})
describe('typeLabel', () => {
// The one bilingual pill, and it has to speak the writer's own pair — a
// hardcoded 翻译 would be simply wrong on the screen of a Portuguese writer.
it('names a translation in the writers own language', () => {
expect(typeLabel('translate', zh)).toBe('翻译 · Translate')
expect(typeLabel('translate', ptPT)).toBe('Tradução · Translate')
expect(typeLabel('translate', fr)).toBe('Traduction · Translate')
})
it('leaves every other type in English, whatever the pair', () => {
for (const p of [zh, ptPT, fr]) {
for (const type of ALL.filter((t) => t !== 'translate')) {
expect(typeLabel(type, p), `${type} on ${p.code}`).toBe(TYPE_META[type].label)
}
}
})
})
@@ -1,4 +1,5 @@
import type { SuggestionType } from '../../api/client'
import type { Pack } from '../../i18n'
// Per-type accent color + human label, mirroring the design tokens. Shared by the
// inline hover SuggestionCard and the margin SuggestionRail so a given suggestion
@@ -8,7 +9,32 @@ export const TYPE_META: Record<SuggestionType, { color: string; label: string }>
phrasing: { color: 'var(--color-peach)', label: 'Phrasing' },
idiom: { color: 'var(--color-lavender)', label: 'Idiom' },
clarity: { color: 'var(--color-sky)', label: 'Clarity' },
// The label here is only the fallback (and what a screen reader gets if the
// pack hasn't arrived yet) — see typeLabel.
translate: { color: 'var(--color-jade)', label: 'Translate' },
voice: { color: 'var(--color-honey)', label: 'Voice' },
collocation: { color: 'var(--color-blossom)', label: 'Word pairing' },
mechanics: { color: 'var(--color-sage)', label: 'Tidy-up' },
}
// typeLabel is what the pill says. It exists for exactly one type: a translation
// card names itself in the writer's own language first, because that language is
// its entire subject. Every other type keeps its English name — those are the
// terms she is learning, and she is learning them in English.
//
// Taking the pack rather than reading the module singleton keeps the label
// reactive: the pair language isn't known until /api/me answers, and a card
// rendered before then must relabel itself when it does.
export function typeLabel(type: SuggestionType, pack: Pack): string {
return type === 'translate' ? pack.editor.translateLabel : TYPE_META[type].label
}
// The label on the accept-a-whole-category button. It stays English even on a
// translation card, whose pill is bilingual: the pill names the kind of advice
// she is reading, while this names an action taken over the rest of the queue,
// and every other word in the action row — Accept, Dismiss, Ask Petal — is
// English too. A control that switched languages between cards would read as a
// different control.
export function batchLabel(type: SuggestionType, count: number): string {
return `Accept all ${TYPE_META[type].label} (${count})`
}
+30 -1
View File
@@ -19,6 +19,9 @@ interface Props {
// True when Petal can't reach its LLM helper — shows a gentle, reassuring note
// (the writing still saves locally, so this is awareness, not an error).
llmDown: boolean
// How many suggestions the rail is holding. Shown as a soft bilingual line
// beside the word count; hidden at zero.
suggestionCount: number
}
// Save-state labels. English except for the lapsed-session case, which is the
@@ -48,9 +51,23 @@ interface Indicator {
label: string
}
export function StatusBar({ wordCount, text, saveStatus, checking, voicing, collocating, llmDown }: Props) {
export function StatusBar({
wordCount,
text,
saveStatus,
checking,
voicing,
collocating,
llmDown,
suggestionCount,
}: Props) {
const t = usePack()
const label = saveLabel(saveStatus, t)
// The gentle counterpart of Grammarly's score: how much is waiting, never how
// well she wrote. Nothing waiting says itself — an empty rail — so a zero here
// would only be a verdict delivered after every check, which is the pressure
// this app's non-goals rule out.
const petals = suggestionCount > 0 ? t.status.petalsToPolish(suggestionCount) : null
const indicators: Indicator[] = [
{
@@ -109,6 +126,18 @@ export function StatusBar({ wordCount, text, saveStatus, checking, voicing, coll
</button>
{statsOpen && <StatsPanel text={text} wordCount={wordCount} />}
</div>
{petals && (
<>
<span aria-hidden>·</span>
{/* Native half first, as everywhere else in Petal. (The review wrote it
English-first; the pack's order is the one she reads all day.) */}
<span className="inline-flex items-center gap-1.5">
<span aria-hidden>🌸</span>
<span>{petals.native}</span>
<span style={{ opacity: 0.75 }}>· {petals.en}</span>
</span>
</>
)}
{indicators
.filter((i) => i.active)
.map((i) => (
Binary file not shown.
+39
View File
@@ -0,0 +1,39 @@
import { useEffect, useState } from 'react'
import { loadSegmenter, type Segmenter } from '../lib/segment'
// Loads the Chinese word list, once per session, and only for a writer who is
// going to use it.
//
// Modelled on useSpellChecker, and gated harder. That hook loads for everyone,
// because everyone's English gets spell-checked; this one loads a megabyte for
// the one direction that needs it, and an account practising English would
// never ask a single question of it. The gate is the writer's own setting rather
// than a guess from their text: a Mandarin native drafting English quotes
// Chinese in it constantly, and none of that is what this is for.
//
// A failure resolves to null, which every consumer already handles as "no
// segmentation" — the Chinese hover quietly does nothing rather than the editor
// refusing to open.
export function useSegmenter(enabled: boolean): Segmenter | null {
const [segmenter, setSegmenter] = useState<Segmenter | null>(null)
useEffect(() => {
if (!enabled) {
// Turning the direction back drops it. It is a megabyte of resident map
// whose only consumer just switched off, and re-loading costs one fetch
// that the browser cache answers.
setSegmenter(null)
return
}
let cancelled = false
loadSegmenter().then((seg) => {
if (!cancelled) setSegmenter(seg)
})
return () => {
cancelled = true
}
}, [enabled])
return segmenter
}
+23 -1
View File
@@ -42,5 +42,27 @@ export function useSession() {
}
}, [])
return { me, signedOut }
// Turn the pair around. The account is the source of truth for which
// direction the editor is in — it decides whether the word list loads at all —
// so the state moves only once the server has agreed, and it moves to what the
// server *stored* rather than to what was asked for.
const setDirection = async (direction: string) => {
const updated = await api.setDirection(direction)
setMe(updated)
}
// Move the pair, naming the direction with it. The two are validated together
// server-side, so an account that is learning Chinese cannot change pair by
// sending `pair_lang` alone — the combination it would ask for (French with
// segmentation) does not exist and is refused. Saying both is how that move is
// made, and routing it through here rather than through the picker's own
// `api` call is what keeps `me.direction` — which decides whether the word
// list stays loaded — in step with what was actually stored.
const setPair = async (lang: string, direction: string) => {
const updated = await api.setPair(lang, direction)
setPackLang(updated.pair_lang)
setMe(updated)
}
return { me, signedOut, setDirection, setPair }
}
+17
View File
@@ -98,6 +98,23 @@ const PAIR_DICTS: Partial<Record<PairLang, DictSpec>> = {
extendedAlphabet: true,
elision: FR_ELISION,
},
// Spanish glues its pronouns onto the *end* of a verb rather than little words
// onto the front, so there is no elision list here — and none is needed: the
// upstream dictionary carries the enclitic forms itself (dámelo, hacérselo,
// escribiéndolo), because unlike French's thirty-four prefix rules they do not
// multiply the word list into the megabytes.
//
// This is RLA's *generic* build, not Debian's hunspell-es — the latter is the
// peninsular one under a pan-Hispanic-looking name, and it rejects vení and
// tenés. See web/public/dictionaries/es/LICENSE: the whole of Spanish is
// accepted here, because underlining is the only thing this file can do.
es: {
lang: 'es',
aff: '/dictionaries/es/es.aff',
dic: '/dictionaries/es/es.dic.gz',
gzipped: true,
extendedAlphabet: true,
},
}
// Where the list lived before it had an owner (Phase 7). Read once, handed to
+178 -6
View File
@@ -4,12 +4,27 @@ import { onPackChange, pack, resetPackForTests, setPackLang, shippedPacks } from
import { zh } from './packs/zh'
import { ptPT } from './packs/pt-PT'
import { fr } from './packs/fr'
import { es } from './packs/es'
import type { Pack } from './types'
// Every pack that ships. Shape assertions run over all of them, because the
// point of Phase 19 was that a language is data — and data that only the first
// author's pack satisfies isn't a shape, it's a coincidence.
const PACKS: Pack[] = [zh, ptPT, fr]
const PACKS: Pack[] = [zh, ptPT, fr, es]
// Every string a pack would ever put on screen, and nothing else — field names
// excluded (see the pt-PT grep below for what including them cost). Templates are
// invoked so an interpolated line is checked as she'd read it, with the same
// stand-in arguments the old JSON replacer used.
function copyOf(p: Pack): string {
const strings = (v: unknown): string[] => {
if (typeof v === 'string') return [v]
if (typeof v === 'function') return strings((v as (...a: unknown[]) => unknown)(1, 'x'))
if (v && typeof v === 'object') return Object.values(v).flatMap(strings)
return []
}
return strings(p).join('\n')
}
beforeEach(() => {
resetPackForTests()
@@ -35,8 +50,10 @@ describe('pack selection', () => {
it('falls back rather than blanking on a pair with no pack yet', () => {
// A pair_lang the deployment has no copy for is a deployment that got ahead
// of its translation. She should still get a working editor. (This was 'fr'
// until Phase 24 gave fr a pack; 'es' is the pair still waiting for one.)
setPackLang('es')
// until Phase 24, then 'es' until Phase 25 — every pair PairLang names now
// has a pack, so the stand-in is a regional code Petal has not decided
// about, which is the realistic version of this failure anyway.)
setPackLang('es-ES')
expect(pack()).toBe(zh)
setPackLang('klingon')
expect(pack()).toBe(zh)
@@ -56,7 +73,7 @@ describe('pack selection', () => {
expect(seen).not.toHaveBeenCalled()
// An unshipped pair resolves back to zh, which is also not a change.
setPackLang('es')
setPackLang('es-ES')
expect(seen).not.toHaveBeenCalled()
setPackLang('pt-PT')
@@ -69,7 +86,7 @@ describe('pack selection', () => {
// matching allowlist exists to enforce from the other side.
it('offers exactly the pairs it has copy for', () => {
const codes = shippedPacks().map((p) => p.code)
expect(codes.sort()).toEqual(['fr', 'pt-PT', 'zh'])
expect(codes.sort()).toEqual(['es', 'fr', 'pt-PT', 'zh'])
// Every offered pair names itself, because a writer stranded on the wrong
// pack can only read the label that is in her own language.
for (const p of shippedPacks()) expect(p.nativeName.length).toBeGreaterThan(0)
@@ -150,6 +167,42 @@ describe('the zh pack', () => {
if (p.code === 'pt-PT') expect(p.locale).toBe('pt-PT') // never pt-BR
})
// The chat-failure line is the only message the Ask Petal panel writes without
// the model, and it renders through the same bilingual bubble as a real reply
// (splitBilingual, blank line between the halves). A pack that writes it as
// one language gets a bubble with a muted empty half — and, worse, tells the
// half of the pair that can't read that language nothing at all.
it.each(PACKS)('says the chat-failure line in both halves of the pair ($code)', async (p) => {
const { splitBilingual } = await import('../components/Editor/bilingualReply')
const { native, en } = splitBilingual(p.editor.chatFailed)
expect(native, `${p.code} chatFailed has no pair-language half`).not.toBe('')
expect(en, `${p.code} chatFailed has no English half`).not.toBe('')
// The English half is the one every reader of every pack shares, so it is
// the one worth pinning: a pack that translated it has lost the point.
expect(en).toMatch(/ask me again/i)
expect(native).not.toBe(en)
})
// The status-bar count is the one line in Petal that grows a number, so it is
// the one that can be quietly ungrammatical in three languages at once — and
// the count itself has to survive translation, since it is the whole content.
it.each(PACKS)('counts petals to polish in both halves, and agrees on the number ($code)', (p) => {
for (const n of [1, 2, 5, 21]) {
const { native, en } = p.status.petalsToPolish(n)
expect(native, `${p.code} has no pair-language half for n=${n}`).toBeTruthy()
expect(en, `${p.code} has no English half for n=${n}`).toBeTruthy()
expect(native, `${p.code} drops the count from its native half`).toContain(String(n))
expect(en).toContain(String(n))
// English is the half every pack shares; a pack that translated it has lost
// the point, exactly as with chatFailed above.
expect(en).toMatch(/petals? to polish/)
}
// One is not many, in every language Petal ships.
expect(p.status.petalsToPolish(1).native).not.toBe(p.status.petalsToPolish(2).native)
expect(p.status.petalsToPolish(1).en).toContain('petal to polish')
expect(p.status.petalsToPolish(2).en).toContain('petals to polish')
})
it.each(PACKS)('labels every companion, tone and style ($code)', async (p) => {
const { COMPANIONS } = await import('../components/Companion/companions')
for (const c of COMPANIONS) {
@@ -194,10 +247,23 @@ describe('the pt-PT pack', () => {
// is invisible to anyone who doesn't read Portuguese — including whoever
// reviews this diff.
it('is European Portuguese, not Brazilian', () => {
// The pack's *copy*, and only its copy. Keys are English identifiers and can
// never be Brazilian, but they used to be in this haystack — and a naive
// substring grep duly failed on `translateLabel`, which lowercases to
// "transla·tela·bel" and so "contains" the pt-BR *tela*. A test that fires on
// its own field names is a test that gets deleted the third time it does it.
//
// Lowercased: half this copy is sentences, and a form that only ever appears
// at the start of one ("Actualmente…") would otherwise walk straight past
// both greps below.
const text = JSON.stringify(ptPT, (_k, v) => (typeof v === 'function' ? v(1, 'x') : v)).toLowerCase()
const text = copyOf(ptPT).toLowerCase()
// Canaries. Every assertion below is a *negative*, so a copyOf that quietly
// returned nothing — or stopped invoking templates — would make this whole
// test pass by having nothing to search. The second line is inside a
// function, and only appears if templates are still being called.
expect(text, 'copyOf collected no pack copy').toContain('sinónimos')
expect(text, 'copyOf stopped invoking templates').toContain('já vais em 1 palavras')
// Brazilian spellings and vocabulary that would give the pack away.
for (const bad of ['sinônimo', 'acadêmico', 'arquivo', 'tela', 'salvar', 'deletar', 'usuário', 'você']) {
@@ -301,6 +367,97 @@ describe('the fr pack', () => {
})
})
describe('the es pack', () => {
// Spanish has no regional question in its dictionary at all — hunspell-es
// ships twenty country codes and every one is a symlink to one pan-Hispanic
// word list — so, even more than with French, the entire regional decision
// lives in this file. Neutral Latin American was chosen deliberately, and a
// stray peninsular form is invisible to everyone reviewing the diff.
it('is Latin American, not peninsular', () => {
const text = copyOf(es).toLowerCase()
for (const bad of [
'ordenador', 'vosotros', 'zumo', 'patata', 'coche',
'gafas', 'billete', 'chaval', 'guay',
// The one that is not merely regional: *coger* is an everyday verb in
// Spain and obscene through most of Latin America. A companion in a
// private notebook must never produce it by accident.
'coger',
]) {
expect(text, `peninsular form "${bad}" in the es pack`).not.toMatch(
new RegExp(`\\b${bad}\\b`),
)
}
// And the accent it is read aloud in. Six of Piper's nine Spanish voices
// are es_ES, so the wrong country is the easy default here — the pt-PT
// trap, not the fr non-question.
expect(es.locale).toBe('es-MX')
})
// Spanish opens its questions and exclamations, and the easiest place to
// forget is exactly where a reviewer's eye slides past: inside a Line's
// native half, and inside an interpolated template. Every opening mark in the
// pack is deliberate; a missing one is a typo the type system cannot see.
it('opens every question and exclamation it closes', () => {
const lines: { native: string; where: string }[] = []
const walk = (node: unknown, path: string) => {
if (typeof node === 'function') return walk((node as (...a: unknown[]) => unknown)(1, 'x'), path)
if (!node || typeof node !== 'object') return
const rec = node as Record<string, unknown>
// A Line is the one shape whose `native` is pure Spanish — the flat
// "Spanish · English" strings carry English punctuation too, so they are
// asserted by hand below rather than by rule.
if (typeof rec.native === 'string' && typeof rec.en === 'string') {
lines.push({ native: rec.native, where: path })
return
}
for (const [k, v] of Object.entries(rec)) walk(v, path ? `${path}.${k}` : k)
}
walk(es, '')
expect(lines.length).toBeGreaterThan(40)
for (const { native, where } of lines) {
if (native.includes('?')) expect(native, `${where} closes ? without ¿`).toContain('¿')
if (native.includes('!')) expect(native, `${where} closes ! without ¡`).toContain('¡')
}
// An exclamative opening with Qué/Cómo/Cuánto is the case that slips past a
// reader, because it carries no closing "!" to look wrong against — the
// whole pair is simply absent. The quorum review caught exactly one of
// these ("Qué linda elección de palabra"), and only one reviewer of four
// saw it, which is the argument for asserting it instead of re-reviewing it.
for (const { native, where } of lines) {
expect(native, `${where} opens an exclamative without ¡`).not.toMatch(
/^(Qué|Cómo|Cuánto|Cuánta)\b/,
)
}
// And the two-language strings, by hand: both halves punctuate their own way.
expect(es.garden.promptRecognition).toBe('¿Qué significa? · What does this mean?')
expect(es.garden.promptProduction).toContain('¿Cuál es la palabra en inglés?')
})
it('renders its interpolated lines with the value in place', () => {
expect(es.app.duplicateTitle('Primavera')).toBe('Primavera (copia)')
expect(es.companion.milestone(300).native).toContain('300 palabras')
// Spanish agreement is the pack's business; the call site only ever passes
// a number. `flor`/`flores` is the irregular one — it takes -es, not -s.
expect(es.garden.reviewDue(1)).toContain('1 palabra ·')
expect(es.garden.reviewDue(4)).toContain('4 palabras ·')
expect(es.garden.growing(1)).toContain('1 flor en el jardín')
expect(es.garden.growing(3)).toContain('3 flores en el jardín')
expect(es.journal.kept(1)).toContain('1 cosa que te llevaste')
expect(es.journal.kept(5)).toContain('5 cosas que te llevaste')
expect(es.status.petalsToPolish(1).native).toContain('1 pétalo por pulir')
expect(es.status.petalsToPolish(2).native).toContain('2 pétalos por pulir')
})
it('says the collision line, which this pair meets constantly', () => {
// real, red, once, pie, sin, pan, mayor, sale, ropa — Spanish and English
// collide about as often as French and English do.
expect(es.editor.alsoIn).toBeTruthy()
expect(es.editor.alsoIn).not.toBe(zh.editor.alsoIn)
expect(es.editor.alsoIn).not.toBe(fr.editor.alsoIn)
expect(es.editor.alsoIn).not.toBe(ptPT.editor.alsoIn)
})
})
// False friends are a per-pair dataset rather than copy: the Latin pairs carry
// the traps their writers actually fall into, and the zh pair legitimately has
// none. Both halves of that are worth pinning.
@@ -328,6 +485,21 @@ describe('false friends', () => {
expect(Object.keys(fr.falseFriends).length).toBeGreaterThan(10)
})
it('the es pair carries the ones that cost most', () => {
// "embarrassed" is the reason this feature exists at all: *embarazada* is
// "pregnant", and it is the single false friend most likely to be said out
// loud to a room. "molest" is the other one that has to be here, because
// *molestar* is an everyday word and the English is not.
for (const word of ['embarrassed', 'molest', 'actually', 'realize', 'exit', 'carpet']) {
expect(es.falseFriends[word], word).toBeDefined()
}
// Spanish shares more Latin with English than either of the other Latin
// pairs, so this list is the longest of the four and should stay that way.
expect(Object.keys(es.falseFriends).length).toBeGreaterThan(
Object.keys(fr.falseFriends).length,
)
})
it('is keyed by the lowercase English word, so a lookup can find it', () => {
for (const p of PACKS) {
for (const key of Object.keys(p.falseFriends)) {
+6 -3
View File
@@ -18,12 +18,15 @@ import type { Pack, PairLang } from './types'
import { zh } from './packs/zh'
import { ptPT } from './packs/pt-PT'
import { fr } from './packs/fr'
import { es } from './packs/es'
export type { Pack, PairLang, Line } from './types'
// Every pack Petal ships. es is the same two lines when its copy is written —
// TypeScript names every string a new pack still owes.
const PACKS: Partial<Record<PairLang, Pack>> = { zh, 'pt-PT': ptPT, fr }
// Every pack Petal ships — and now every pair PairLang names, so this map is
// no longer Partial by necessity. It stays Partial anyway: the next pair will
// be declared in the type before its copy exists, exactly as es was, and the
// gap between the two is the point.
const PACKS: Partial<Record<PairLang, Pack>> = { zh, 'pt-PT': ptPT, fr, es }
const DEFAULT_LANG: PairLang = 'zh'
+550
View File
@@ -0,0 +1,550 @@
// The Spanish pack — the fourth pair, and the second written straight into the
// groove Phase 24 cut.
//
// ⚠️ REVIEWED BY FOUR MODELS, NOT BY A NATIVE SPEAKER.
// SUGGESTIONS.md §3 sets the bar: a pack should be reviewed by someone who
// speaks the pair before it is trusted. That has still not happened. What has
// happened (2026-07-28) is the same interim pass the fr and pt-PT packs got —
// four models read this file independently as Latin American Spanish speakers,
// and only findings at least two of them reached on their own were applied,
// listed in BUILD_PLAN Phase 25. A quorum of models agreeing is agreement, not
// authority: it can catch a verb form no one says and a register that slips into
// Spain, and it cannot catch a line that is correct and lifeless. Treat this as a
// better-checked draft.
//
// The choices this file makes, and why:
//
// * **Latin American neutral, not peninsular.** Chosen deliberately (user,
// 2026-07-28) over es-ES: it is the Spanish far more people write, and the
// "neutral" register is a real thing that Spanish-language publishing and
// dubbing have spent decades stabilising. So: *tú* for the singular,
// **ustedes** for the plural and no *vosotros* anywhere, and the pan-American
// half of every vocabulary split — *computadora*, *celular*, *carro*, *jugo*,
// *papa*, *departamento*, *lentes*, *boleto*. A vitest greps this file for the
// peninsular twins the way the fr pack is grepped for québécismes, because a
// stray *ordenador* is invisible to everyone reviewing the diff.
// * **The dictionary makes the opposite choice on purpose, and that is not a
// contradiction.** This copy is Latin American; the spelling dictionary
// behind it accepts *every* variety of Spanish, peninsular and voseante
// alike (RLA's generic build — see web/public/dictionaries/es/LICENSE). The
// two answer different questions: the copy is Petal *speaking*, where a
// register has to be chosen, and the dictionary is Petal *listening*, where
// the only available action is to underline something. Choosing a register
// to write in costs a reader nothing; choosing one to accept would tell her
// that her own conjugation is a typo.
// * **Tuteo.** Petal is a companion in someone's private notebook, and *usted*
// would put a desk between them — the same call the pt-PT pack made about
// *tu* over *você* and the fr pack about *tu* over *vous*.
// * **Tuteo here, voseo accepted there.** This copy says *tú*, because neutral
// Latin American is tuteo and something had to be chosen. The dictionary
// nonetheless accepts *vení* and *tenés*, so a Rioplatense writer is never
// told her own present tense is a misspelling — she simply reads a companion
// that speaks a slightly different Spanish than she writes, which is true of
// every Spanish speaker reading anything.
// * **¿Inverted marks, siempre!** Spanish opens questions and exclamations and
// this file does too — including inside interpolated lines, where it is
// easiest to forget. It is also the one punctuation habit that costs her
// nothing in English: unlike the French space before « ! », there is no
// Spanish mark to carry across by accident, so no prose note has to warn
// about it.
// * Quotation marks are the curly “ ” rather than « », which is the American
// convention and the commoner one in Spanish outside Spain.
//
// Spanish first, English underneath — same shape as the other packs, for the same
// reason: she reads her own language faster, and the English half is what she is
// here to learn.
import type { Pack } from '../types'
export const es: Pack = {
code: 'es',
nativeName: 'Español',
// Mexican Spanish is the standard "neutral" broadcast variety and the one
// Piper has a Latin American voice for (es_MX-ald-medium). The peninsular
// es_ES-davefx-medium the plan originally named would have read this copy in
// the accent it was written to avoid.
locale: 'es-MX',
app: {
duplicateTitle: (title) => `${title} (copia)`,
garden: 'Jardín de palabras',
history: 'Historial',
},
auth: {
title: 'Vuelve a entrar',
titleEn: 'Please sign in again',
bodyWithDraft:
'Lo que acabas de escribir está guardado en este dispositivo — vuelve a entrar y se guardará solo.',
bodyWithDraftEn: "What you just wrote is safe on this device — it'll save itself once you're back in.",
bodyPlain: 'Tu sesión expiró. Todo lo que escribiste ya está guardado.',
bodyPlainEn: 'Your session expired. Everything you wrote is already saved.',
signIn: 'Iniciar sesión · Sign in',
},
companion: {
choose: 'Elige un compañero · Choose a companion',
encouragements: [
{ native: '¡Eso! Esa oración fluye mucho mejor 🌸', en: 'Lovely — that reads so much smoother now.' },
{ native: 'Cada vez escribes mejor ✨', en: "You're getting better and better." },
{ native: 'Me gusta mucho ese cambio 💕', en: 'I really like that change.' },
{ native: '¡Sigue así, lo estás logrando!', en: 'Keep going — youve got this!' },
{ native: 'Mmm, así queda mucho más claro 👍', en: 'Mm, thats much clearer.' },
{ native: '¡Qué linda elección de palabra! 🌷', en: 'Thats such a good word choice.' },
{ native: 'Ay, ese párrafo se lee solito ☁️', en: 'Ooh, that paragraph flows so nicely.' },
{ native: 'Me encanta verte escribir con más confianza 💛', en: 'I love watching you write with more confidence.' },
{ native: 'Cada avance cuenta, por chiquito que sea 🌱', en: 'Every little bit of progress counts.' },
{ native: 'Hoy tus palabras están brillando ✨', en: 'Your words are sparkling today.' },
],
tips: [
{ native: 'Consejo: en inglés, las oraciones cortas se leen mejor.', en: 'Tip: shorter English sentences often read clearer.' },
{ native: 'No olvides los artículos “the” y “a”.', en: "Don't forget articles like “the” and “a”." },
{ native: 'Para el pasado, usa el pretérito: go → went.', en: 'For the past, use past tense: go → went.' },
{ native: 'Leer en voz alta ayuda a notar lo que suena raro.', en: 'Reading aloud helps you catch awkward spots.' },
{ native: 'Una idea por párrafo y todo queda clarito.', en: 'One idea per paragraph keeps it tidy.' },
{ native: '¿Tienes una duda? Pregúntame ✨', en: 'Not sure about something? Just ask me. ✨' },
{ native: 'El plural lleva “s”: two apples 🍎', en: 'Plurals take an “s”: two apples 🍎' },
// Two tips the zh pack has no use for. Spanish and English share so much
// Latin vocabulary that the false friends are a daily hazard, and the
// subject pronoun is the habit Spanish speakers drop most often.
{ native: 'Cuidado con los falsos amigos: “actually” no significa *actualmente*.', en: 'Careful with false friends — “actually” means *in fact*.' },
{ native: 'En inglés el sujeto casi nunca se omite: “it is raining”, no “is raining”.', en: 'English almost always needs a subject: “it is raining”, not “is raining”.' },
],
breaks: [
{ native: 'Llevas rato escribiendo — estírate y descansa la vista 🍵', en: "You've been writing a while — stretch and rest your eyes. 🍵" },
{ native: '¿Un vaso de agua y cinco minutos de pausa?', en: 'Sip some water and take five?' },
{ native: 'Mira a lo lejos un momento, la vista te lo agradece 🌿', en: 'Look into the distance for a moment — give your eyes a break. 🌿' },
],
// Late-night nudges. The English wit is the user's own and is kept word for
// word across every pack; the Spanish line leads gently into it, exactly as
// the Mandarin, Portuguese and French ones do.
bedtime: [
{ native: 'Tu cama ha de estar preguntándose dónde andas 🛏️', en: 'I bet your bed is missing you right now.' },
{ native: 'Cansada, escribes mal — ve a descansar 🌙', en: 'A tired writer is a bad writer — get some rest.' },
{ native: 'Mejor consúltalo con la almohada ✨', en: 'Sleep is a wondrous enabler.' },
{ native: '¿Oyes? No… porque todos están dormidos, y tú deberías estarlo también 😴', en: "Hear that? No… you don't, because everyone is sleeping and you should be too." },
// Spanish sayings about sleep and haste, in place of the French and
// Portuguese ones — a pack is not a translation of another pack.
{ native: 'Dormir es el mejor remedio.', en: 'Sleep is the best medicine.' },
{ native: 'A quien madruga, Dios lo ayuda.', en: 'The early riser gets a hand from above.' },
{ native: 'No por mucho madrugar amanece más temprano.', en: 'Rising earlier will not make the sun come up sooner.' },
],
greeting: { native: '¡Hola! Aquí te hago compañía 🐱', en: "Hi! I'm right here keeping you company. 🐱" },
welcomeBack: { native: '¡Volviste! ✨ Sigamos', en: 'Welcome back ✨ lets keep going!' },
errors: [
{ native: 'Uy — un tropiezo chiquito, pero tus palabras están a salvo.', en: 'Oops — a little hiccup, but your words are safe.' },
{ native: 'Ay, me enredé un segundo — ya vuelvo.', en: 'Haiya, I got stuck for a sec — back in a moment.' },
{ native: 'No te preocupes, lo intentamos de nuevo en un ratito 🍵', en: "Don't worry — let's try again in a bit. 🍵" },
],
milestone: (words: number) => ({
native: `¡Guau! Ya llevas ${words} palabras 🎉`,
en: `Wow — ${words} words already! Amazing. 🎉`,
}),
// Algo pequeño que escribir, ofrecido una vez al día a una página en blanco.
// Recuerdos y opiniones, nunca ejercicios: aquí no hay nada que se pueda
// reprobar, y esa es justamente la intención.
invitations: [
{ native: 'Escribe 50 palabras: algo pequeño que te hizo sonreír hoy 🌸', en: 'Write 50 words: one small thing that made you smile today.' },
{ native: 'Escribe 50 palabras: lo más rico que comiste hoy', en: 'Write 50 words: the best thing you ate today.' },
{ native: 'Escribe 50 palabras: lo que ves por tu ventana en este momento', en: 'Write 50 words: what you can see out of your window right now.' },
{ native: 'Escribe 50 palabras: un lugar al que volverías con gusto', en: 'Write 50 words: somewhere you would happily go back to.' },
{ native: 'Escribe 50 palabras: algo que aprendiste esta semana', en: 'Write 50 words: one thing you learned this week.' },
{ native: 'Escribe 50 palabras: un mensaje para ti misma dentro de un año', en: 'Write 50 words: something to tell yourself a year from now.' },
{ native: 'Escribe 50 palabras: una canción que no has parado de escuchar estos días', en: 'Write 50 words: a song you have had on lately.' },
{ native: 'Escribe 50 palabras: alguien a quien te gustaría agradecerle hoy', en: 'Write 50 words: someone you would like to thank today.' },
],
inviteAccept: 'Vamos · Lets write',
inviteDecline: 'Hoy no · Not today',
declined: { native: 'Está bien, me regreso a dormir 😴', en: 'Fair enough — back to my nap. 😴' },
names: {
cat: 'Gato dormilón',
dog: 'Perro alegre',
'wiggle-dog': 'Perro meneacola',
butterfly: 'Mariposa',
parrot: 'Loro',
},
},
prose: {
longSentence: 'Esta oración quedó un poco larga — partirla en dos o tres la hace más clara 🌸',
commaSplice: 'Aquí hay dos oraciones unidas solo por una coma. Pon un punto, o únelas con “and / but”.',
vagueThis: (word) => `No queda claro a qué se refiere “${word}” — precísalo (por ejemplo “${word} idea / change…”).`,
oxfordComma: 'En una lista de tres o más elementos, una coma antes de “and / or” ayuda a leer (la coma de Oxford).',
transitionComma: (word) => `Después de un conector al inicio de la oración va una coma: “${word}, …”.`,
capitalizeSentence: 'Empieza cada oración con mayúscula.',
repeatedWord: (word) => `Parece que “${word}” quedó escrito dos veces — échale un ojo.`,
capitalizeI: 'En inglés, “I” (yo) siempre va con mayúscula.',
spaceBeforePunct: 'En inglés no se deja espacio antes de la puntuación: la coma y el punto van pegados a la palabra, igual que en español.',
spaceAfterPunct: 'Después de una coma o un punto, deja un espacio antes de la siguiente palabra.',
articleAn: (word) => `Antes de un sonido de vocal se usa “an”: “an ${word}”.`,
articleA: (word) => `Antes de un sonido de consonante se usa “a”: “a ${word}”.`,
uncountable: (word, singular) => `${word}” es incontable en inglés — no lleva s, basta con “${singular}”.`,
capitalizeProper: (fixed) => `En inglés, los idiomas, las nacionalidades, los días y los meses van con mayúscula: “${fixed}”.`,
thirdPersonS: (subject, verb) => `Con he/she/it, el verbo lleva -s: “${subject} ${verb}”.`,
pluralAfter: (determiner, noun) => `Después de “${determiner}”, el sustantivo va en plural: “${determiner} ${noun}s”.`,
doubleDeterminer: (first, second) => `${first} ${second}” lleva dos determinantes — deja solo uno (quita “${first}”, por ejemplo).`,
thereArePlural: (noun) => `En plural se dice “there are”: “there are ${noun}…”.`,
itsOwn: '“its” = “it is”. Para decir “su”, es “its” — o sea “its own”.',
itsIs: (rest) => `Aquí va “its ${rest}” (it is); “its” es el posesivo.`,
thanNotThen: (word) => `En una comparación se escribe “than”, no “then”: “${word} than”.`,
preposition: (wrong, right) => `En inglés se dice “${right}”, no “${wrong}” — esa preposición es fija.`,
collocation: (wrong, right) => `En inglés estas palabras van juntas así: “${right}”, y no “${wrong}”.`,
doubleComparative: (lead, word) => `${word}” ya es el comparativo — no necesita “${lead}”: con “${word}” basta.`,
peopleArePlural: (verb) => `“People” es plural en inglés: “people ${verb}”.`,
// Interferencia del español. Las tres primeras son las que de verdad se ven
// todos los días; el resto las tiene el pack aunque este par no las corra.
ageIsNotHave: (years) => `En inglés la edad se dice con *to be*, no con *tener*: “I am ${years} years old”.`,
agreeIsAVerb: '“Agree” ya es el verbo — no lleva *to be* delante: se dice “I agree”, no “I am agree”.',
forNotSince: (duration) => `Para una duración se usa “for”: “for ${duration}”. “Since” marca el punto de partida (since 2020).`,
veryBeforeVerb: (verb) => `“Very” solo acompaña adjetivos, no verbos: “really ${verb}”, o “${verb}… very much”.`,
turnOnNotOpen: (thing, on) => `En inglés los aparatos no se abren, se encienden: “turn ${on ? 'on' : 'off'} the ${thing}”.`,
althoughOrBut: (word) => `En inglés se pone “${word}” o “but”, nunca los dos en la misma oración.`,
},
// Los falsos amigos entre el español y el inglés — la trampa que hace sentir
// ridícula en vez de simplemente corregida. Por eso son solo un aviso: Petal
// nunca cambia la palabra, porque “actually” bien pudo ser la que quería.
//
// El español comparte tanto latín con el inglés que esta lista es la más larga
// de los cuatro packs, y *embarrassed* es la razón por la que existe la función.
falseFriends: {
embarrassed: {
native: '“Embarrassed” significa *apenada, avergonzada*. *Embarazada* se dice “pregnant”.',
en: '“Embarrassed” means ashamed; the Spanish *embarazada* is “pregnant”.',
},
actually: {
native: '“Actually” significa *en realidad*, no *actualmente*. Para *actualmente* se dice “currently” o “nowadays”.',
en: '“Actually” means *in fact*. For the Spanish *actualmente*, English uses “currently”.',
},
eventually: {
native: '“Eventually” significa *al final, tarde o temprano* — no *eventualmente*. Para eso: “possibly” o “if necessary”.',
en: '“Eventually” means *in the end*, not *possibly*.',
},
realize: {
native: '“Realize” significa *darse cuenta*. Para *realizar* (llevar a cabo) se dice “carry out” o “do”.',
en: '“Realize” means to become aware; *realizar* is “to carry out”.',
},
assist: {
native: '“Assist” significa *ayudar*. Para *asistir* (ir a algo) se dice “attend”.',
en: '“Assist” means to help; *asistir a* is “to attend”.',
},
attend: {
native: '“Attend” significa *asistir a*. Para *atender* (a alguien) se dice “serve” o “take care of”.',
en: '“Attend” means to go to something; *atender* is “to serve”.',
},
support: {
native: '“Support” significa *apoyar*. Para *soportar* (aguantar algo molesto) se dice “put up with” o “bear”.',
en: '“Support” means to back someone up; *soportar* is “to put up with”.',
},
sensible: {
native: '“Sensible” significa *sensato*. Para *sensible* se dice “sensitive”.',
en: '“Sensible” means level-headed; the Spanish *sensible* is “sensitive”.',
},
sympathetic: {
native: '“Sympathetic” significa *comprensivo, solidario*. Para *simpático* se dice “nice” o “friendly”.',
en: '“Sympathetic” means understanding; *simpático* is “nice”.',
},
library: {
native: '“Library” es la *biblioteca*. La *librería* se dice “bookshop” o “bookstore”.',
en: '“Library” is where books are lent; a shop that sells them is a “bookstore”.',
},
exit: {
native: '“Exit” es la *salida*. El *éxito* se dice “success”.',
en: '“Exit” is the way out; *éxito* is “success”.',
},
success: {
native: '“Success” es el *éxito*. Un *suceso* se dice “event”.',
en: '“Success” is achievement; a Spanish *suceso* is an “event”.',
},
carpet: {
native: '“Carpet” es la *alfombra*. Una *carpeta* se dice “folder”.',
en: '“Carpet” covers a floor; a *carpeta* is a “folder”.',
},
discussion: {
native: '“Discussion” es una *conversación*, sin pelea. Una *discusión* (riña) se dice “argument”.',
en: '“Discussion” is calm; a Spanish *discusión* is an “argument”.',
},
argument: {
native: '“Argument” es una *discusión* o *pelea*. El *argumento* de una historia se dice “plot”.',
en: '“Argument” is a quarrel; the *argumento* of a story is its “plot”.',
},
introduce: {
native: '“Introduce” es *presentar* a alguien. Para *introducir* (meter) se dice “insert” o “put in”.',
en: '“Introduce” is to present someone; *introducir* is “to insert”.',
},
molest: {
native: '“Molest” significa *abusar sexualmente* — nunca se usa por *molestar*. Para eso: “bother” o “annoy”.',
en: '“Molest” means to abuse; the everyday *molestar* is “to bother”.',
},
constipated: {
native: '“Constipated” significa *estreñida*. Para *constipada* (resfriada) se dice “to have a cold”.',
en: '“Constipated” is a bowel problem; *constipado* is “a cold”.',
},
large: {
native: '“Large” significa *grande*. Para *largo* se dice “long”.',
en: '“Large” means big; *largo* is “long”.',
},
lecture: {
native: '“Lecture” es una *conferencia* o *clase*. La *lectura* se dice “reading”.',
en: '“Lecture” is a talk; *lectura* is “reading”.',
},
parents: {
native: '“Parents” son los *padres*. Los *parientes* se dicen “relatives”.',
en: '“Parents” are your mother and father; *parientes* are “relatives”.',
},
record: {
native: '“Record” significa *grabar* o *registrar*. Para *recordar* se dice “remember”.',
en: '“Record” means to register; *recordar* is “to remember”.',
},
remove: {
native: '“Remove” significa *quitar*. Para *remover* (revolver) se dice “stir”.',
en: '“Remove” means to take away; *remover* is “to stir”.',
},
rope: {
native: '“Rope” es la *cuerda*. La *ropa* se dice “clothes”.',
en: '“Rope” is cord; *ropa* is “clothes”.',
},
once: {
native: '“Once” significa *una vez*. El número *once* se dice “eleven”.',
en: '“Once” means one time; the Spanish *once* is “eleven”.',
},
question: {
native: '“Question” es una *pregunta*. Una *cuestión* (asunto) se dice “matter” o “issue”.',
en: '“Question” is something you ask; a *cuestión* is a “matter”.',
},
compromise: {
native: '“Compromise” es un *acuerdo con concesiones*. Un *compromiso* (obligación) se dice “commitment”.',
en: '“Compromise” is meeting halfway; a *compromiso* is a “commitment”.',
},
career: {
native: '“Career” es la *trayectoria profesional*. La *carrera* que se estudia se dice “degree” o “major”.',
en: '“Career” is your working life; a university *carrera* is a “degree”.',
},
ultimately: {
native: '“Ultimately” significa *a fin de cuentas*. Para *últimamente* se dice “lately”.',
en: '“Ultimately” means in the end; *últimamente* is “lately”.',
},
idiom: {
native: '“Idiom” es una *expresión hecha*. El *idioma* se dice “language”.',
en: '“Idiom” is a set phrase; *idioma* is “language”.',
},
deception: {
native: '“Deception” significa *engaño*. La *decepción* se dice “disappointment”.',
en: '“Deception” means being misled; *decepción* is “disappointment”.',
},
pretend: {
native: '“Pretend” significa *fingir*. Para *pretender* (aspirar a) se dice “intend” o “claim”.',
en: '“Pretend” means to fake; *pretender* is “to intend”.',
},
},
docs: {
sortRecent: 'Recientes · Recent',
sortTitle: 'Título · Title',
sortLongest: 'Los más largos · Longest',
backUpAll: 'Respaldar todo · Back up all:',
signOut: 'Cerrar sesión · Sign out',
duplicate: 'Duplicar · Duplicate',
searchPlaceholder: 'Buscar · Search',
searching: 'Buscando… · Searching…',
noMatches: 'Sin resultados · No matches',
tags: 'Etiquetas · Tags',
newTagPlaceholder: 'Nueva etiqueta · New tag',
language: 'Idioma · Language',
languageFailed: 'No se pudo cambiar el idioma — sigue igual · Couldnt switch',
},
editor: {
askPlaceholder: 'Ask why… / Pregunta por qué…',
chatFailed: 'No logré responder — vuelve a preguntarme, por favor. 🌸\n\nI had trouble answering just now — please ask me again. 🌸',
findPlaceholder: 'Buscar · Find',
findNone: 'Nada · 0',
matchCase: 'Match case · Distinguir mayúsculas',
close: 'Close · Cerrar',
replacePlaceholder: 'Reemplazar por · Replace',
replace: 'Reemplazar',
replaceAll: 'Todo',
translateLabel: 'Traducción · Translate',
spelling: 'Ortografía · Spelling',
noSuggestions: 'Sin sugerencias · No suggestions',
addToDictionary: 'Agregar al diccionario · Add to dictionary',
readSelection: 'Leer la selección en voz alta · Read selection aloud',
rewrite: 'Reescribir · Rewrite',
rewriting: 'Reescribiendo… · Rewriting…',
rewriteFailed: 'No se pudo reescribir — inténtalo otra vez · Couldnt rewrite',
cancel: 'Cancelar · Cancel',
retry: 'Reintentar · Retry',
useThis: 'Usar esta · Use this',
word: 'Palabra · Word',
inGarden: 'Ya está en el jardín · In your garden (tap to remove)',
saveToGarden: 'Guardar en el jardín · Save to garden',
readAloud: 'Leer en voz alta · Read aloud',
readSlowly: 'Leer despacio · Read slowly',
readAloudNative: 'Leer en español · Read in Spanish',
lookingUp: 'Buscando… · Looking up…',
definition: 'Definición · Definition',
synonyms: 'Sinónimos · Synonyms',
tapToSwap: 'toca para cambiar · tap to swap',
nothingFound: 'No encontré esta palabra · Nothing found for this word',
origin: 'Origen · Origin',
// El par español ve esto seguido: inglés y español comparten tanto latín que
// los choques son la regla y no la excepción — real, red, once, pie, sin,
// pan, mayor, sale, ropa.
alsoIn: 'También es una palabra en español · Also a word in Spanish',
wordBands: {
simple: { native: 'De todos los días', en: 'Everyday word' },
standard: { native: 'Común', en: 'Standard' },
advanced: { native: 'Avanzada', en: 'Advanced' },
},
},
styles: {
natural: { native: 'Más natural', en: 'Natural' },
academic: { native: 'Académico', en: 'Academic' },
professional: { native: 'Profesional', en: 'Professional' },
casual: { native: 'Informal', en: 'Casual' },
humorous: { native: 'Divertido', en: 'Humorous' },
creative: { native: 'Creativo', en: 'Creative' },
persuasive: { native: 'Persuasivo', en: 'Persuasive' },
},
tones: {
general: { native: 'General', en: 'General' },
academic: { native: 'Académico', en: 'Academic' },
professional: { native: 'Profesional', en: 'Professional' },
casual: { native: 'Informal', en: 'Casual' },
humorous: { native: 'Divertido', en: 'Humorous' },
creative: { native: 'Creativo', en: 'Creative' },
persuasive: { native: 'Persuasivo', en: 'Persuasive' },
},
exports: {
label: 'Exportar',
print: 'Imprimir / PDF',
formats: {
md: { native: 'Markdown', en: 'Markdown (.md)' },
docx: { native: 'Documento de Word', en: 'Word (.docx)' },
html: { native: 'Página web', en: 'Web page (.html)' },
txt: { native: 'Texto sin formato', en: 'Plain text (.txt)' },
},
},
garden: {
title: 'Jardín de palabras · Vocabulary Garden',
titleWithFlower: '🌷 Jardín de palabras · Vocabulary Garden',
reviewing: 'Repaso · Reviewing — recall, then grade yourself',
subtitle: 'Words you looked up, blooming as you learn them',
reviewDue: (n) => `Repasar ${n} palabra${n === 1 ? '' : 's'} · Review ${n} due 🌸`,
emptyLead: 'Tu jardín todavía está vacío.',
emptyHint: 'Haz clic derecho en una palabra en inglés para buscarla — y aquí brotará.',
due: 'por repasar · due',
seen: (reps, intervalDays) => `repasada ${reps}× · seen ${reps}× · intervalo ${intervalDays} d`,
readAloud: '🔊 Leer',
readSlowly: '🐢 Despacio',
source: '📄 Fuente',
remove: '🗑 Quitar',
growing: (n) => `🐱💤 ${n} flor${n === 1 ? '' : 'es'} en el jardín · ${n} blossom${n > 1 ? 's' : ''} growing`,
end: 'Terminar · End',
promptProduction: '¿Cuál es la palabra en inglés? · Which English word?',
promptRecognition: '¿Qué significa? · What does this mean?',
showAnswer: 'Ver la respuesta · Show answer',
gradeAgain: { native: 'Repetir', en: 'Again' },
gradeGood: { native: 'La recuerdo', en: 'Good' },
gradeEasy: { native: 'Fácil', en: 'Easy' },
},
journal: {
tabGarden: '🌷 Jardín · Garden',
tabJournal: '🌱 Avances · Growth',
subtitle: 'Your own writing, month by month — only ever you and your past self',
empty: 'Sigue escribiendo un poco más — esta página crece a partir de tu propio trabajo. · Keep writing; this page grows out of your own work.',
keptHead: 'Este mes · This month',
kept: (n) => `${n} cosa${n === 1 ? '' : 's'} que te llevaste · ${n} thing${n === 1 ? '' : 's'} you took on board`,
keptBefore: (n) => `${n} el mes pasado · ${n} the month before`,
stuckHead: 'Se te quedó · Stayed with you',
stuck: (phrase, docs) =>
`${phrase}” — ya la escribes sola, en ${docs} de tus textos · now in ${docs} of your pieces`,
fadedHead: 'Ya no necesitas corregir esto · You stopped needing this',
faded: (pattern, times) =>
`${pattern}” — ${times}× antes, ninguna este mes · ${times}× back then, none this month`,
cheerStuck: (phrase) => ({
native: `¡Ya usas “${phrase}” tú solita! 🌱`,
en: `Youre using “${phrase}” on your own now! 🌱`,
}),
cheerFaded: (pattern) => ({
native: `Hace rato que “${pattern}” no necesita corrección 😌`,
en: `${pattern}” hasnt needed fixing in a while 😌`,
}),
},
history: {
title: 'Historial · History',
kinds: {
manual: { native: 'Punto guardado', en: 'Saved point' },
auto: { native: 'Automático', en: 'Auto' },
pre_restore: { native: 'Antes de restaurar', en: 'Before restore' },
},
justNow: 'ahora mismo · just now',
minutesAgo: (n) => `hace ${n} min · ${n} min ago`,
hoursAgo: (n) => `hace ${n} h · ${n} hr ago`,
daysAgo: (n) => `hace ${n} día${n > 1 ? 's' : ''} · ${n} day${n > 1 ? 's' : ''} ago`,
preview: 'Vista previa · Preview',
restoring: 'Restaurando… · Restoring…',
restoreThis: 'Restaurar esta versión · Restore this version',
passport: '📜 Pasaporte de escritura · Writing passport',
keepFullHistory: 'Guardar todo el historial · Keep full history',
},
status: {
savedLocally: 'Guardado en este dispositivo · Kept on this device',
helperRestingNative: 'El ayudante está descansando',
helperRestingEn: "· Petal's helper is resting · tu texto está guardado",
petalsToPolish: (n) => ({
native: `${n} ${n === 1 ? 'pétalo' : 'pétalos'} por pulir`,
en: `${n} ${n === 1 ? 'petal' : 'petals'} to polish`,
}),
soundsOn: 'Sonidos activados · Sounds on',
soundsOff: 'Sonidos desactivados · Sounds off',
petalsOn: 'Pétalos activados · Petals on',
petalsOff: 'Pétalos desactivados · Petals off',
statsTitle: 'Estadísticas · Writing stats',
stats: {
words: { native: 'Palabras', en: 'Words' },
characters: { native: 'Caracteres', en: 'Characters' },
sentences: { native: 'Oraciones', en: 'Sentences' },
paragraphs: { native: 'Párrafos', en: 'Paragraphs' },
pages: { native: 'Páginas', en: 'Pages' },
readingTime: { native: 'Tiempo de lectura', en: 'Reading time' },
avgWordLength: { native: 'Longitud promedio', en: 'Avg word length' },
variety: { native: 'Variedad de vocabulario', en: 'Word variety' },
readability: { native: 'Nivel de lectura', en: 'Reading level' },
},
readability: {
easy: { native: 'Fácil', en: 'Easy' },
standard: { native: 'Común', en: 'Standard' },
fairlyHard: { native: 'Algo difícil', en: 'Fairly hard' },
advanced: { native: 'Avanzado', en: 'Advanced' },
},
},
toolbar: {
untitledHeading: '(sin título)',
outline: 'Esquema · Outline',
outlineHint: 'Usa H1/H2/H3 para crear títulos, y el esquema aparecerá aquí.',
},
update: {
available: 'Hay una versión nueva disponible',
refresh: 'Actualizar · Refresh',
dismiss: 'Más tarde · Dismiss',
},
}
+6
View File
@@ -307,6 +307,7 @@ export const fr: Pack = {
editor: {
askPlaceholder: 'Ask why… / Demande pourquoi…',
chatFailed: 'Je nai pas réussi à répondre — repose-moi la question, sil te plaît. 🌸\n\nI had trouble answering just now — please ask me again. 🌸',
findPlaceholder: 'Rechercher · Find',
findNone: 'Rien · 0',
matchCase: 'Match case · Respecter la casse',
@@ -314,6 +315,7 @@ export const fr: Pack = {
replacePlaceholder: 'Remplacer par · Replace',
replace: 'Remplacer',
replaceAll: 'Tout',
translateLabel: 'Traduction · Translate',
spelling: 'Orthographe · Spelling',
noSuggestions: 'Aucune suggestion · No suggestions',
addToDictionary: 'Ajouter au dictionnaire · Add to dictionary',
@@ -448,6 +450,10 @@ export const fr: Pack = {
savedLocally: 'Gardé sur cet appareil · Kept on this device',
helperRestingNative: 'Lassistant se repose',
helperRestingEn: "· Petal's helper is resting · ton texte est enregistré",
petalsToPolish: (n) => ({
native: `${n} ${n === 1 ? 'pétale' : 'pétales'} à polir`,
en: `${n} ${n === 1 ? 'petal' : 'petals'} to polish`,
}),
soundsOn: 'Sons activés · Sounds on',
soundsOff: 'Sons coupés · Sounds off',
petalsOn: 'Pétales activés · Petals on',
+6
View File
@@ -285,6 +285,7 @@ export const ptPT: Pack = {
editor: {
askPlaceholder: 'Ask why… / Pergunta porquê…',
chatFailed: 'Não consegui responder agora — pergunta-me outra vez, se faz favor. 🌸\n\nI had trouble answering just now — please ask me again. 🌸',
findPlaceholder: 'Localizar · Find',
findNone: 'Nada · 0',
matchCase: 'Match case · Maiúsculas/minúsculas',
@@ -292,6 +293,7 @@ export const ptPT: Pack = {
replacePlaceholder: 'Substituir por · Replace',
replace: 'Substituir',
replaceAll: 'Tudo',
translateLabel: 'Tradução · Translate',
spelling: 'Ortografia · Spelling',
noSuggestions: 'Sem sugestões · No suggestions',
addToDictionary: 'Adicionar ao dicionário · Add to dictionary',
@@ -425,6 +427,10 @@ export const ptPT: Pack = {
savedLocally: 'Guardado neste dispositivo · Kept on this device',
helperRestingNative: 'O ajudante está a descansar',
helperRestingEn: "· Petal's helper is resting · o teu texto está guardado",
petalsToPolish: (n) => ({
native: `${n} ${n === 1 ? 'pétala' : 'pétalas'} para polir`,
en: `${n} ${n === 1 ? 'petal' : 'petals'} to polish`,
}),
soundsOn: 'Som ligado · Sounds on',
soundsOff: 'Som desligado · Sounds off',
petalsOn: 'Pétalas ligadas · Petals on',
+19
View File
@@ -15,6 +15,17 @@ export const zh: Pack = {
nativeName: '中文',
locale: 'zh-CN',
// The zh pair is the only one Petal can be *learned* toward, because it is the
// only one with a word list and a Chinese→English dictionary (Phase 26). The
// two labels are each written for the person who would pick them: she reads
// the first, and the English speaker learning her language reads the second.
learner: {
label: '我在学 · I am learning',
toEn: '英文',
toPair: 'Chinese 中文',
failed: '没能换成功 · Couldnt switch — nothing changed',
},
app: {
duplicateTitle: (title) => `${title} (副本)`,
garden: '词汇花园',
@@ -185,6 +196,7 @@ export const zh: Pack = {
editor: {
askPlaceholder: 'Ask why… / 问为什么…',
chatFailed: '我这会儿没答上来,再问我一次好吗?🌸\n\nI had trouble answering just now — please ask me again. 🌸',
findPlaceholder: '查找 · Find',
findNone: '无 · 0',
matchCase: 'Match case · 区分大小写',
@@ -192,6 +204,7 @@ export const zh: Pack = {
replacePlaceholder: '替换为 · Replace',
replace: '替换',
replaceAll: '全部',
translateLabel: '翻译 · Translate',
spelling: '拼写 · Spelling',
noSuggestions: '没有建议 · No suggestions',
addToDictionary: '添加到词典 · Add to dictionary',
@@ -325,6 +338,12 @@ export const zh: Pack = {
savedLocally: '已保存在本机 · Kept on this device',
helperRestingNative: '小助手在休息',
helperRestingEn: "· Petal's helper is resting · 文字已保存",
// 片 is the measure word for petals; a digit reads perfectly naturally in
// Chinese and spares every pack a numeral table.
petalsToPolish: (n) => ({
native: `${n}片花瓣待打磨`,
en: `${n} ${n === 1 ? 'petal' : 'petals'} to polish`,
}),
soundsOn: '声音开 · Sounds on',
soundsOff: '声音关 · Sounds off',
petalsOn: '花瓣开 · Petals on',
+38
View File
@@ -39,6 +39,29 @@ export interface Pack {
// Portuguese voice anyone reaches for is Brazilian.
locale: string
// Copy for turning this pair around — a writer who is native in English and
// learning X, rather than the other way round.
//
// Optional, and its presence is the pack's half of the same fact
// auth.learnerPairs holds server-side: a pair can only be learned toward if
// Petal has a word list to segment it with and a dictionary that reads from it
// into English. Chinese has both; the Latin pairs have neither yet, so their
// packs simply leave this out and the control does not render.
//
// Each label is written in the language of the person who would *choose* it,
// for the same reason the pair buttons name themselves: someone on the wrong
// side of this switch cannot read the side they are trying to reach.
learner?: {
// The heading over the two choices.
label: string
// "I am practising English" — read by the writer who is native in X.
toEn: string
// "I am learning X" — read by the writer who is native in English.
toPair: string
// Shown when the server refuses the change.
failed: string
}
app: {
// A duplicated document's title. A function, not a suffix: where the marker
// goes is the pack's business.
@@ -155,6 +178,10 @@ export interface Pack {
editor: {
askPlaceholder: string
// Shown in Petal's own chat bubble when the reply never arrives. It is
// the one line in that panel Petal writes without the model, so the pack
// owns it — and it is bilingual like every answer beside it.
chatFailed: string
findPlaceholder: string
findNone: string
matchCase: string
@@ -162,6 +189,12 @@ export interface Pack {
replacePlaceholder: string
replace: string
replaceAll: string
// The one suggestion-type pill that is bilingual. Every other type name
// (Grammar, Phrasing, Idiom) stays English: they are the vocabulary of the
// thing she is learning, and she is learning it in English. This card is the
// opposite case — its whole subject is her own language — so it says so in
// her language first. The pack holds the rendered string, separator and all.
translateLabel: string
spelling: string
noSuggestions: string
addToDictionary: string
@@ -287,6 +320,11 @@ export interface Pack {
savedLocally: string
helperRestingNative: string
helperRestingEn: string
// How many suggestions are waiting, said gently — the status-bar counterpart
// of the rail. Never a score: it counts what is there, and says nothing about
// how well she is writing. Only ever called with n >= 1 (nothing waiting is
// said by the rail being empty, not by a badge announcing it).
petalsToPolish: (n: number) => Line
soundsOn: string
soundsOff: string
petalsOn: string
+32 -2
View File
@@ -22,6 +22,12 @@
--color-honey: #CE9B4F; /* voice */
--color-blossom: #E59ABF; /* collocation — warm blossom pink */
--color-sage: #B7C7B9; /* mechanics — calm sage, a gentle tidy-up nudge */
/* translate jade. Deliberately the one saturated teal in the set: this is the
card that carries her own language across into English, and it should be
findable at a glance in a stack of pastels. Jade rather than a cooler blue
because the colour is warmly auspicious in the language it most often
serves, and it stays clear of grammar's mint and clarity's sky. */
--color-jade: #6FB5B8;
--color-success: #8FCFA8; /* saved / accepted */
/* Typography */
@@ -233,6 +239,7 @@ button, a, input {
.petal-suggestion-phrasing { border-bottom-color: var(--color-peach); }
.petal-suggestion-idiom { border-bottom-color: var(--color-lavender); }
.petal-suggestion-clarity { border-bottom-color: var(--color-sky); }
.petal-suggestion-translate { border-bottom-color: var(--color-jade); }
.petal-suggestion-voice { border-bottom-color: var(--color-honey); }
.petal-suggestion-collocation { border-bottom-color: var(--color-blossom); }
.petal-suggestion-mechanics { border-bottom-color: var(--color-sage); }
@@ -307,6 +314,19 @@ button, a, input {
color: var(--color-plum);
}
/* Accept-all: the quieter sibling of Accept, on both card surfaces. It's outlined
in its category's own colour rather than filled like Accept, because it acts on
cards she can't see from here an equally loud button for a larger action would
invite the click she meant to give the one suggestion in front of her. */
.petal-accept-all {
border: 1px solid;
background: transparent;
transition: background-color 120ms ease;
}
.petal-accept-all:hover {
background: var(--color-surface-alt);
}
/* --- Find & Replace ---------------------------------------------------------
In-document search (Ctrl/Cmd+F). Every match gets a soft honey wash; the
current match is brighter with a rose ring so it stands out as you step
@@ -405,9 +425,19 @@ button, a, input {
.petal-companion {
/* Mascot size scales with the viewport width: ~original on a laptop, up to
~2× on a large desktop. Tune the middle (vw) term to taste. */
--petal-companion-size: clamp(10rem, 17vw, 20rem);
--petal-companion-size: clamp(9rem, 15.3vw, 18rem);
animation: petal-bob 3.2s ease-in-out infinite;
transition: transform 200ms ease;
/* Shrink toward its corner when fading out of a card's way. `scale` is a
separate property from `transform` so it composes with the bob keyframes. */
transform-origin: bottom right;
transition: transform 200ms ease, opacity 400ms ease, scale 400ms ease;
}
/* A suggestion card has drifted into the corner: the kitten politely turns
translucent, steps back 10%, and lets clicks pass through to the card. */
.petal-companion-faded {
opacity: 0.15;
scale: 0.9;
pointer-events: none;
}
/* The Lottie art sits inside the round badge with a little breathing room. */
.petal-companion-art {
+19
View File
@@ -0,0 +1,19 @@
// fromIME answers whether a keydown belongs to an in-flight IME composition
// rather than to the app.
//
// While a candidate window is open, Enter and Escape mean something to the IME
// and nothing to Petal: Enter commits the candidate, Escape cancels it back to
// the pinyin. A handler that acts on them anyway steals the key — she presses
// Escape to fix a wrong candidate and the sidebar reappears; she presses Enter
// to accept 公园 and the Find bar jumps to the next match instead. In neither
// case does the IME get its keystroke.
//
// `isComposing` is the standard signal and is what modern browsers set. The 229
// keyCode is the older one, still the only signal some Safari/IME combinations
// give, and costs one comparison to honour.
// React's synthetic keyboard event doesn't surface `isComposing`, so the native
// event underneath it is what gets asked — the same object either way.
export function fromIME(e: KeyboardEvent | { nativeEvent: KeyboardEvent }): boolean {
const native = 'nativeEvent' in e ? e.nativeEvent : e
return native.isComposing || native.keyCode === 229
}
+219
View File
@@ -0,0 +1,219 @@
import { readFileSync } from 'node:fs'
import { gunzipSync } from 'node:zlib'
import { describe, expect, it } from 'vitest'
import { buildSegmenter, isHan, type Segmenter } from './segment'
// Segmentation is tested twice over, and the two halves check different things.
//
// The hand-built dictionaries below pin the *algorithm*: given these words with
// these frequencies, this is the split, and the reason is visible in the four
// lines above the assertion. They would pass with any word list.
//
// The block at the bottom pins the *shipped asset*: the real 188,522-word list
// this app serves, on the sentences a rebuild would plausibly break. Those are
// the cases where being wrong is invisible — the app still works, it just
// underlines and glosses the wrong thing.
// A dictionary written the way the asset is: "word freq" per line.
function dict(entries: Record<string, number>): Segmenter {
return buildSegmenter(
Object.entries(entries)
.map(([w, f]) => `${w} ${f}`)
.join('\n'),
)
}
const words = (seg: Segmenter, text: string) => seg.segment(text).map((t) => t.word)
describe('isHan', () => {
it('accepts Han across the extension blocks, and nothing else', () => {
expect(isHan('中')).toBe(true)
expect(isHan('龥')).toBe(true)
// Beyond the basic block. A character Petal fails to recognise as Chinese is
// one the English tokenizer then tries to make sense of.
expect(isHan('𠀀')).toBe(true)
for (const ch of ['a', '1', ' ', '', '。', 'あ', '한']) {
expect(isHan(ch), ch).toBe(false)
}
})
})
describe('the walk chooses the likeliest split, not the longest match', () => {
// The textbook case, and the reason longest-match is not good enough: 研究生
// ("graduate student") is a real word and a longer match than 研究 at position
// 0 — but 研究/生命 ("research" + "life") is the likelier path, and it is the
// sentence a person would read.
it('研究生命的起源', () => {
const seg = dict({ 研究: 6000, 研究生: 800, 生命: 4000, : 900, : 300000, 起源: 700 })
expect(words(seg, '研究生命的起源')).toEqual(['研究', '生命', '的', '起源'])
})
it('乒乓球拍卖完了 — the ambiguity is 球拍 against 拍卖', () => {
const seg = dict({
乒乓球: 500, 乒乓: 400, 球拍: 200, 拍卖: 900, 卖完: 50, : 3000, : 200000, : 2000, : 800,
})
expect(words(seg, '乒乓球拍卖完了')).toEqual(['乒乓球', '拍卖', '完', '了'])
})
it('keeps particles as their own words', () => {
const seg = dict({ : 90000, : 300000, 中文: 3000, : 20000, : 60000, : 40000, : 50000 })
expect(words(seg, '他的中文说得很好')).toEqual(['他', '的', '中文', '说', '得', '很', '好'])
})
})
describe('what the walk does with what it does not know', () => {
// A sentence with an unfamiliar character in it must still segment. Every
// position needs *some* path through it, which is why an unknown character
// scores badly rather than not scoring at all.
it('an unknown character becomes its own token and the rest survives', () => {
const seg = dict({ : 90000, 喜欢: 5000, : 2000 })
expect(words(seg, '我喜欢龥猫')).toEqual(['我', '喜欢', '龥', '猫'])
})
// It must never *invent* a word: an unknown span of two characters is two
// unknown characters, not a new headword.
it('never joins unknown characters into a word', () => {
const seg = dict({ : 90000 })
expect(words(seg, '我龥龥')).toEqual(['我', '龥', '龥'])
})
// A character above the BMP is two UTF-16 code units, and asking about either
// half alone says "not Han". Getting this wrong is quiet: the run breaks in
// two around the character, the words either side of it stop being looked up,
// and nothing anywhere reports an error.
it('a supplementary-plane character is one unknown token inside the run', () => {
const seg = dict({ : 90000, 喜欢: 5000, : 2000 })
const text = '我喜欢𠀀猫'
expect(words(seg, text)).toEqual(['我', '喜欢', '𠀀', '猫'])
for (const t of seg.segment(text)) expect(text.slice(t.from, t.to)).toBe(t.word)
// And it is hoverable from either code unit — a caret offset can land on
// the low surrogate, which is not a character boundary but is a real index.
expect(seg.wordAt(text, 3)?.word).toBe('𠀀')
expect(seg.wordAt(text, 4)?.word).toBe('𠀀')
expect(seg.wordAt(text, 5)?.word).toBe('猫')
})
// A rare real word still loses to two common ones — this is the property that
// lets the shipped list keep 100,000 rare CC-CEDICT headwords without them
// distorting ordinary sentences.
it('a rare long word loses to two common short ones', () => {
const seg = dict({ 公园: 4000, 跑步: 3000, 公园跑: 1 })
expect(words(seg, '公园跑步')).toEqual(['公园', '跑步'])
})
})
describe('Chinese is not the only thing in the paragraph', () => {
const seg = dict({ : 90000, : 50000, : 8000, 英文: 3000 })
// Latin runs are skipped, not returned. The English tokenizer is still running
// over the same text and owns them; returning them here would mean two layers
// claiming one word.
it('skips Latin and punctuation, keeping offsets into the original string', () => {
const tokens = seg.segment('我在写 English 英文。')
expect(tokens.map((t) => t.word)).toEqual(['我', '在', '写', '英文'])
for (const t of tokens) {
expect('我在写 English 英文。'.slice(t.from, t.to)).toBe(t.word)
}
})
it('每 token reports the span it actually occupies', () => {
const tokens = seg.segment('英文')
expect(tokens).toEqual([{ word: '英文', from: 0, to: 2 }])
})
})
describe('wordAt — the hover and click path', () => {
const seg = dict({ : 90000, 今天: 8000, : 30000, 公园: 4000, 跑步: 3000, : 200000 })
const text = '我今天去公园跑步了'
it('finds the word covering a position anywhere inside it', () => {
// 公园 occupies [4,6): either of its characters resolves to the whole word
// rather than to one character.
for (const i of [4, 5]) {
expect(seg.wordAt(text, i)?.word, `index ${i}`).toBe('公园')
}
expect(seg.wordAt(text, 0)?.word).toBe('我')
expect(seg.wordAt(text, 2)?.word).toBe('今天')
})
// A position names a gap; a word covers characters. On a boundary the answer
// is the word that *starts* there, because that is the character being pointed
// at — index 6 is the 跑 under the mouse, not the 园 behind it.
it('a boundary belongs to the word that starts there', () => {
expect(seg.wordAt(text, 6)?.word).toBe('跑步')
expect(seg.wordAt(text, 4)?.word).toBe('公园')
})
// The caret after a just-typed word belongs to that word. Ctrl/Cmd+D at the
// end of 跑步 must look up 跑步, which is the position the caret is actually in
// the moment someone finishes typing it.
it('a caret at the very end of the text still resolves', () => {
expect(seg.wordAt(text, text.length)?.word).toBe('了')
})
it('returns null outside Han text', () => {
expect(seg.wordAt('hello world', 3)).toBeNull()
expect(seg.wordAt('', 0)).toBeNull()
expect(seg.wordAt('我 hello', 4)).toBeNull()
})
// The window exists so that a pasted page of Chinese with no punctuation is
// not walked on every hover. It must not change the answer for ordinary text.
it('agrees with a full segmentation of the same string', () => {
const long = '我今天去公园跑步了'.repeat(20)
const full = seg.segment(long)
for (const t of full) {
expect(seg.wordAt(long, t.from)).toEqual(t)
}
})
})
// ── the shipped asset ───────────────────────────────────────────────────────
// Everything above would pass with a word list built wrong. These read the file
// this app actually serves.
describe('the shipped word list', () => {
const raw = gunzipSync(readFileSync(new URL('../../public/dictionaries/zh/words.txt.gz', import.meta.url)))
const seg = buildSegmenter(raw.toString('utf8'))
it('is the size the build script says it is', () => {
expect(seg.size).toBeGreaterThan(180_000)
})
it('segments ordinary learner prose the way a reader would', () => {
expect(words(seg, '我今天早上去公园跑步了')).toEqual(['我', '今天', '早上', '去', '公园', '跑步', '了'])
expect(words(seg, '他的中文说得很好')).toEqual(['他', '的', '中文', '说', '得', '很', '好'])
expect(words(seg, '北京大学的学生正在图书馆学习')).toEqual([
'北京大学', '的', '学生', '正在', '图书馆', '学习',
])
})
it('gets the textbook ambiguities right', () => {
expect(words(seg, '研究生命的起源')).toEqual(['研究', '生命', '的', '起源'])
expect(words(seg, '乒乓球拍卖完了')).toEqual(['乒乓球', '拍卖', '完', '了'])
})
// The minimal pair, and the one that says the line above was a decision rather
// than a bias against long words: the same five characters open both
// sentences, and 研究生 is the right answer in one of them.
it('finds 研究生 where 研究生 is the word', () => {
expect(words(seg, '研究生宿舍')).toEqual(['研究生', '宿舍'])
expect(words(seg, '他们正在研究生物')).toEqual(['他们', '正在', '研究', '生物'])
})
// The three particles the 错别字 rules are about have to survive as their own
// tokens, or those rules have nothing to anchor to.
it('keeps 的 / 地 / 得 separate', () => {
expect(words(seg, '她高兴地笑了')).toContain('地')
expect(words(seg, '这个问题需要认真地思考')).toContain('地')
expect(words(seg, '他跑得很快')).toContain('得')
expect(words(seg, '我的书')).toContain('的')
})
it('knows the words the build script asserts it kept', () => {
for (const w of ['我', '的', '图书馆', '乒乓球', '公园', '的士']) {
expect(seg.has(w), w).toBe(true)
}
})
})
+271
View File
@@ -0,0 +1,271 @@
// Chinese word segmentation — the thing that has to exist before any of Petal's
// ESL surfaces can point at a Chinese word.
//
// Every one of them is built on `wordAt(doc, pos)`, and `wordAt` is a regex over
// runs of Latin letters. That works because English writes its word boundaries
// down. Chinese does not: 我今天早上去公园跑步了 is eleven characters and seven
// words, and which seven is a question with a real answer that no regex can
// reach. Until something answers it there is no "word under the cursor" to
// hover, look up, read aloud, or plant in the vocabulary garden.
//
// **Why the answer is a shortest-path walk and not longest-match.** The obvious
// algorithm — take the longest dictionary word at each position and move on —
// gets the textbook cases wrong in both directions, because the longest match is
// not the likeliest one. The standard fix is to score every possible split by
// how probable its words are and take the best-scoring path, which is a
// shortest-path problem over a small DAG and is what this does. It is why the
// word list ships with a frequency column at all.
//
// **Why it runs in the browser.** It runs on hover. A round-trip per hover is
// not a hover, and the whole point of the offline lexicon (SUGGESTIONS §6) is
// that the daily reading aids keep working with the tunnel down.
// A word found in the text, with the offsets it occupies. Offsets are into the
// string that was passed in — the caller maps them to ProseMirror positions the
// same way the spell and suggestion layers already do.
export interface Token {
word: string
from: number
to: number
}
// Han characters only. Not a hand-rolled U+4E00U+9FFF range: that misses the
// extension blocks, and a character Petal fails to recognise as Chinese is one
// the English tokenizer then tries to make sense of.
const HAN = /^\p{Script=Han}$/u
// `ch` is one character, but "one character" is a code point, not a UTF-16 code
// unit: the extension blocks live above the BMP and `text[i]` there is half a
// surrogate pair. Testing a lone surrogate against \p{Script=Han} says no —
// which would silently undo the whole reason this is a property escape — so the
// pair is joined back up before it is asked about. Anchored, so a two-code-unit
// string has to *be* one Han character rather than merely contain one.
export function isHan(ch: string): boolean {
return HAN.test(ch)
}
// charAt is isHan's companion for scanning a string: it returns the whole code
// point beginning at `i`, so a surrogate pair is asked about as one character.
function charAt(text: string, i: number): string {
const code = text.codePointAt(i)
return code === undefined ? '' : String.fromCodePoint(code)
}
// isHanAt reports whether the code point *beginning* at `i` is Han. A low
// surrogate (the second half of a pair) is never a start, so it answers for the
// pair it belongs to instead — which keeps a run contiguous across it.
export function isHanAt(text: string, i: number): boolean {
const code = text.charCodeAt(i)
if (code >= 0xdc00 && code <= 0xdfff && i > 0) return isHanAt(text, i - 1)
return isHan(charAt(text, i))
}
// How many code units the character beginning at `i` occupies: two for a
// surrogate pair, one for everything else. Every step through a string here goes
// through this, so a supplementary-plane character is never cut in half.
function charLen(text: string, i: number): number {
const code = text.charCodeAt(i)
return code >= 0xd800 && code <= 0xdbff && i + 1 < text.length ? 2 : 1
}
// Where the character *before* `i` begins, or -1 when there is none.
function prevCharStart(text: string, i: number): number {
if (i <= 0) return -1
const j = i - 1
const code = text.charCodeAt(j)
return code >= 0xdc00 && code <= 0xdfff && j > 0 ? j - 1 : j
}
// The longest word the walk will consider at any position. The dictionary
// contains longer entries (chengyu, place names, a few titles), but the cost of
// the walk is linear in this number and the entries beyond it are rare enough
// that paying for them on every hover is the wrong trade. Six characters covers
// every ordinary word and every four-character idiom.
const MAX_WORD_LEN = 6
// What an unknown single character is worth, as a fraction of one occurrence.
// It must be *positive* — every position needs some path through it, or a
// sentence containing one unfamiliar character would have no segmentation at
// all — and it must be small enough that a real one-character word always wins.
// Half an occurrence is below the rarest thing in the list (which is 1) and
// above zero, which is the whole specification.
const UNKNOWN_WEIGHT = 0.5
export interface Segmenter {
// segment splits a whole string. Runs of non-Han text are skipped rather than
// returned: this is the Chinese tokenizer, and the Latin one is still running
// over the same paragraph.
segment(text: string): Token[]
// wordAt returns the token covering `index`, or null when that position is
// not inside Han text. This is the hover/click path, and it segments only the
// run around the position rather than the whole document.
wordAt(text: string, index: number): Token | null
// has reports whether a word is in the list — the 错别字 rules ask, to check
// that a correction they are about to propose is a real word.
has(word: string): boolean
size: number
}
// buildSegmenter turns the raw `word freq` list into something that can answer
// questions about it. Exported for tests, which build tiny dictionaries by hand;
// the app reaches it through loadSegmenter.
export function buildSegmenter(source: string): Segmenter {
const freq = new Map<string, number>()
let total = 0
for (const line of source.split('\n')) {
if (!line) continue
const sp = line.lastIndexOf(' ')
if (sp <= 0) continue
const word = line.slice(0, sp)
const n = Number(line.slice(sp + 1))
if (!Number.isFinite(n) || n <= 0) continue
freq.set(word, n)
total += n
}
// A dictionary with nothing in it would make every log() below -Infinity.
const logTotal = Math.log(Math.max(total, 1))
const unknownScore = Math.log(UNKNOWN_WEIGHT) - logTotal
// The walk, over one run of Han characters.
//
// `best[i]` is the score of the best segmentation of run[i..], and `next[i]`
// is where that segmentation's first word ends. Filling it right-to-left means
// each position only ever reads answers that are already final, which is what
// makes this linear rather than exponential in the number of possible splits.
function walk(run: string, base: number, out: Token[]): void {
const n = run.length
const best = new Float64Array(n + 1)
const next = new Int32Array(n + 1)
best[n] = 0
for (let i = n - 1; i >= 0; i--) {
// Positions inside a surrogate pair are not character boundaries, so no
// path ever arrives at one and nothing below would ever read the answer.
if (i > 0 && charLen(run, i - 1) === 2) continue
let bestScore = -Infinity
let bestEnd = i + charLen(run, i)
// `len` counts *characters*, which is what MAX_WORD_LEN is in and what the
// dictionary is keyed by; `j` counts code units, which is what a slice is
// in. The two differ exactly where a supplementary character sits.
let j = bestEnd
for (let len = 1; len <= MAX_WORD_LEN && j <= n; len++) {
const f = freq.get(run.slice(i, j))
let score: number
if (f === undefined) {
// Only a single unknown character is a candidate. Allowing unknown
// multi-character spans would let the walk invent words.
if (len > 1) {
if (j >= n) break
j += charLen(run, j)
continue
}
score = unknownScore
} else {
score = Math.log(f) - logTotal
}
score += best[j]
if (score > bestScore) {
bestScore = score
bestEnd = j
}
if (j >= n) break
j += charLen(run, j)
}
best[i] = bestScore
next[i] = bestEnd
}
for (let i = 0; i < n; ) {
const end = next[i]
out.push({ word: run.slice(i, end), from: base + i, to: base + end })
i = end
}
}
function segment(text: string): Token[] {
const out: Token[] = []
let i = 0
while (i < text.length) {
if (!isHanAt(text, i)) {
i += charLen(text, i)
continue
}
let j = i
while (j < text.length && isHanAt(text, j)) j += charLen(text, j)
walk(text.slice(i, j), i, out)
i = j
}
return out
}
// How much context a hover segments. The run around the cursor is bounded
// because a pasted page of Chinese with no punctuation is one run, and a hover
// must not walk it. Segmentation is local enough that a window this size
// reaches the same answer as the whole paragraph would: the walk's decisions
// are dominated by the two or three characters either side, and a word longer
// than MAX_WORD_LEN cannot span the window's edge anyway.
const WINDOW = 60
function wordAt(text: string, index: number): Token | null {
if (index < 0 || index > text.length) return null
// `index` names a gap between characters; a word covers characters. So the
// question is resolved on the character at `index` — the one to the *right*
// of the caret — and a boundary belongs to the word that starts there rather
// than the one that ends there. For a hover that is simply correct: index 6
// of 我今天去公园跑步了 is the 跑 being pointed at.
//
// The step back covers the case where there is no character to the right:
// the caret at the end of the text, or against following punctuation. That
// is where the caret sits the instant an IME commits a word, and Ctrl/Cmd+D
// there must look up the word just typed.
let probe = index
if (probe >= text.length || !isHanAt(text, probe)) {
const prev = prevCharStart(text, probe)
if (prev >= 0 && isHanAt(text, prev)) probe = prev
else return null
}
// Never leave the probe inside a surrogate pair: the slice below starts
// there, and half a character is not a character.
if (probe > 0 && charLen(text, probe - 1) === 2) probe -= 1
let start = probe
while (start > 0 && probe - start < WINDOW) {
const prev = prevCharStart(text, start)
if (prev < 0 || !isHanAt(text, prev)) break
start = prev
}
let end = probe
while (end < text.length && isHanAt(text, end) && end - probe < WINDOW) end += charLen(text, end)
const tokens: Token[] = []
walk(text.slice(start, end), start, tokens)
for (const t of tokens) {
if (probe >= t.from && probe < t.to) return t
}
return null
}
return { segment, wordAt, has: (w) => freq.has(w), size: freq.size }
}
// Where the word list lives. Gzipped, like every dictionary Petal ships that is
// bigger than English's.
const WORDS_URL = '/dictionaries/zh/words.txt.gz'
// loadSegmenter fetches and builds the segmenter. One per session, like the
// spelling dictionaries — the cost is the parse, not the download, and paying it
// per document would be paying it per document for no reason.
//
// A failure resolves to null rather than throwing. Petal without segmentation is
// Petal with no Chinese hover, which is a diminished editor; Petal that refused
// to open because a static asset 404ed is no editor at all.
export async function loadSegmenter(url = WORDS_URL): Promise<Segmenter | null> {
try {
const res = await fetch(url)
if (!res.ok || !res.body) return null
const stream = res.body.pipeThrough(new DecompressionStream('gzip'))
const text = await new Response(stream).text()
const seg = buildSegmenter(text)
return seg.size > 0 ? seg : null
} catch {
return null
}
}
+73
View File
@@ -0,0 +1,73 @@
import { describe, expect, it } from 'vitest'
import { SettledSpans, normalizeForDedup } from './settled'
describe('normalizeForDedup', () => {
// These cases are the client half of a pair: every one of them is also asserted
// against the Go implementation in TestNormalizeMatchesTheClient. The server
// normalizes the spans it sends and the client normalizes the findings it
// checks against them, so a divergence would be silent — a dismissed card
// quietly coming back. Add a case to both or neither.
it('folds every quote variant onto one character', () => {
expect(normalizeForDedup('She said “hello”')).toBe("She said 'hello'")
expect(normalizeForDedup('She said "hello"')).toBe("She said 'hello'")
expect(normalizeForDedup('its')).toBe("it's")
expect(normalizeForDedup('its')).toBe("it's")
expect(normalizeForDedup('`code´')).toBe("'code'")
})
it('collapses runs of whitespace and trims', () => {
expect(normalizeForDedup(' a apple\n here ')).toBe('a apple here')
expect(normalizeForDedup('a\tapple')).toBe('a apple')
})
it('is empty for whitespace only, so it can never settle everything', () => {
expect(normalizeForDedup(' \n ')).toBe('')
})
it('leaves her Chinese alone', () => {
expect(normalizeForDedup('我想说这句话')).toBe('我想说这句话')
})
})
describe('SettledSpans', () => {
it('recognises a span it was told about, in any spelling', () => {
const settled = new SettledSpans()
settled.add('a apple')
expect(settled.has('a apple')).toBe(true)
expect(settled.has('a apple')).toBe(true)
expect(settled.has('an apple')).toBe(false)
})
it('takes the document record and later dismissals through the same door', () => {
const settled = new SettledSpans()
settled.add('the the', 'a apple') // as loaded on open
settled.add('two apple') // as dismissed just now
expect(settled.has('the the')).toBe(true)
expect(settled.has('two apple')).toBe(true)
})
// The load is async and she can dismiss a card while it is in flight. Adding
// rather than assigning is what keeps that dismissal — an assignment here would
// hand the card straight back, which is the bug this whole file is about.
it('keeps a dismissal made before the document record lands', () => {
const settled = new SettledSpans()
settled.add('two apple') // she dismisses
settled.add('the the') // the fetch lands afterwards
expect(settled.has('two apple')).toBe(true)
expect(settled.has('the the')).toBe(true)
})
it('forgets everything on reset, because the record is per document', () => {
const settled = new SettledSpans()
settled.add('a apple')
settled.reset()
expect(settled.has('a apple')).toBe(false)
})
it('ignores an empty span rather than storing one nothing can match', () => {
const settled = new SettledSpans()
settled.add(' ')
expect(settled.has('')).toBe(false)
expect(settled.has('anything')).toBe(false)
})
})
+57
View File
@@ -0,0 +1,57 @@
// Whether a span is one she has already settled — accepted or dismissed.
//
// The rule pack (Companion/prose.ts) detects from the document text alone and has
// no memory between runs, so something has to tell it "she has already answered
// this one". The server keeps that record and applies it to everything it stores
// (see buildSuppressor in internal/suggestions/handlers.go); this is the same
// question asked locally, for the 250 ms pass that renders before — and, offline,
// instead of — the server's reply.
//
// Keyed on the original alone, deliberately, because that is how the server keys
// it: dismissing an edit settles the span, not one particular rewrite of it.
// normalizeForDedup mirrors the Go function of the same name: quote variants
// folded onto one character, runs of whitespace collapsed. The editor rewrites
// quotes as she types and a reflowed paragraph changes its line breaks, so
// comparing raw text would miss spans that are plainly the same.
//
// The two implementations must agree, because the server normalizes the strings
// it sends and the client normalizes the findings it compares against them. They
// differ only on characters neither her writing nor the model produces (JS folds
// a BOM as whitespace, Go folds U+0085); a mismatch there costs one redundant
// card, not a wrong one.
export function normalizeForDedup(s: string): string {
return s.replace(/[‘’‚‛“”„″"`´]/g, "'").trim().split(/\s+/).join(' ')
}
// SettledSpans answers "has she already dealt with this?" for a rule-pack
// finding. Built from the server's record when the document opens, and added to
// as she accepts or dismisses, so the answer is right for cards the server has
// never heard of — the provisional ones.
export class SettledSpans {
private spans = new Set<string>()
// Forget everything (on switching documents — the record is per document).
reset(): void {
this.spans = new Set()
}
// Remember a span she has just actioned. Called for every card that leaves the
// rail, not only the provisional ones: a persisted card is suppressed by the
// server, but its reply arrives a round-trip after the local pass has already
// put the card back on screen.
//
// The document's stored record is loaded through this same door rather than by
// replacing the set, so a card she dismisses while that fetch is in flight isn't
// forgotten when it lands.
add(...originals: string[]): void {
for (const original of originals) {
const norm = normalizeForDedup(original)
if (norm) this.spans.add(norm)
}
}
has(original: string): boolean {
return this.spans.has(normalizeForDedup(original))
}
}