From 15398eab4dbd74d5b0c07da2e5158154e9934f87 Mon Sep 17 00:00:00 2001 From: prosolis <5590409+prosolis@users.noreply.github.com> Date: Wed, 29 Jul 2026 00:38:42 -0700 Subject: [PATCH] What the browser found that the tests could not Four defects behind passing tests, and the reason they passed: the fixtures were written by the same hand as the code and all argued their own case. Claude-Session: https://claude.ai/code/session_01GJHNvirh7Hzhc9RL3HAvz7 --- BUILD_PLAN.md | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/BUILD_PLAN.md b/BUILD_PLAN.md index 69c74ff..1d5a9d4 100644 --- a/BUILD_PLAN.md +++ b/BUILD_PLAN.md @@ -443,7 +443,7 @@ Today every correction and every explanation comes back in English, whatever she - Tests worth writing first: doc-level detection per pair (monolingual, 80/20, 50/50, quotation-heavy English), asserting the hysteresis band **from both directions**; the flipped prompt names the target language and drops the ESL framing while the English-document prompt stays **byte-identical** to today's (the path every existing user is on); the two language arguments proven independent by the only pair that can exercise it — a `learning_pair` zh account writing Chinese wants Chinese corrections explained in English, a `learning_en` zh account writing Chinese wants both in Chinese; a handler test in the shape of `pairlang_test.go`; and a language flip invalidating checked chunks. - [x] **Resolved 2026-07-28 (user's call), and the two questions had different answers.** The **vocabulary garden** tags every card with the language of the document it was met in, and reviews all of them. Filtering the queue to the half she is learning was the alternative and was rejected for the writer this is for: the words she met while writing Portuguese are still words she met, and a garden that quietly drops them stops being a record of her reading. Gating harvesting to English documents was rejected harder — a looked-up word would simply vanish. **Read-aloud** follows the document, which needed the verdict to reach the client at all: it was server-only until now. - **Order:** (a) detection + salt + checkpoint/voice prompts — the whole visible win, independently shippable; (b) translate-card direction, Ask Petal seed, `/translate` direction; (c) garden and read-aloud, once the two questions above are answered. -- **(a), (b) and (c) are built — Phase 28 is complete in code.** **Not seen in a browser and not deployed**: the flipped translate card, the skipped seed, and the garden's tagging are asserted through the real `/check`, `/translate` and `/vocab` paths in Go tests, not watched. A rebuild whenever the user wants it, along with Phases 24–27. +- **(a), (b) and (c) are built, deployed, and seen in a browser 2026-07-29 — see the session log.** Four defects only a browser could have found, three of them in code that had passing tests. **Not seen in a browser and not deployed**: the flipped translate card, the skipped seed, and the garden's tagging are asserted through the real `/check`, `/translate` and `/vocab` paths in Go tests, not watched. A rebuild whenever the user wants it, along with Phases 24–27. **Step (a) as built, 2026-07-28** — items 1–4. - `internal/llm/target.go` — `Target{Correct, Explain, Pair}` + `English` (a `Lang` the `langs` map has no business holding: that map answers "which half is hers"). `EnglishTarget(pair)` is the pre-phase behaviour named, and `Flipped()` is the one question the prompts ask. **Three fields, not two**: the collocation coach's parenthetical gloss is addressed to *her* and not to the document, so it reads `Pair` — collapsing it into `Explain` would have silently moved that gloss into English on every English document, which is every document today. @@ -483,6 +483,14 @@ Today every correction and every explanation comes back in English, whatever she - [x] **Phase 14 — companion warmth + bedtime nag + night mode**: more encouraging phrases, a gentle "go to bed" nudge after 11pm, and a calm dark theme + falling stars at night. ✅ (see Phase 14 above) ## Session log +- 2026-07-29: **Phase 28 deployed to petal.parodia.dev and watched, on a pt-PT account** (user: "deploy it and let's see all of this in a browser"). Merged 76dede8..4660550 to main, pre-migration backup taken, 0017/0018 applied, six containers healthy. **What the phase promised, Qwen delivered on the first try**: "dois pão" → "dois pães" explained in Portuguese and naming *European* Portuguese, a phrasing card, an idiom card, and no English rendering anywhere — the "never translate it into English" line held, which was the untestable part. **Then four defects, three of them in code with passing tests.** + 1. **The verdict never reached the editor.** `useAutoSave` discarded the save response, so read-aloud used the value the document had when it was *opened*. My own note claiming the auto-save refreshed it was simply wrong. The hook now hands the row back and App lifts one field. + 2. **Even fixed, one save behind is the wrong cadence.** The pass that decides the verdict runs *after* a save, so the client learned it only on the next keystroke — and read-aloud is reached for precisely when she has stopped typing. Heard twice in the browser before it was believed. `/check`, `/voice` and `/collocation` now answer with `X-Petal-Doc-Lang`; a header, because all three answer with a bare array that every caller reads as one. + 3. **The marker lists were curated against English so tightly they had been curated against writing.** "Esta manhã acordei cedo e fui correr ao longo da marginal." — two marker hits, **zero** English hits, verdict English. The list had no contractions (ao, à, num), none of the tenses a diary is written in (estava, havia, fomos), and none of the words that join two clauses (até, depois, então, onde). Every one of them clears the list's own bar. The floor stays at 3; three is now reachable by prose rather than only by a paragraph that argues its own case. fr and es got the same additions by analogy — neither has an account to catch it live, which is exactly how this survived. Two regression tests, pointed opposite ways. + 4. **A word of her own language could never enter the garden at all.** Right-clicking "carro" showed a full card — gloss, phonetic, two definitions — and stored nothing: the forward lookup of a Portuguese word answers empty and everything rendered comes from `reverse`, which the capture gate never read. Academic while every document was English; on a Portuguese document it emptied the garden of exactly the words she met. So the tagging shipped that morning was correct and **unreachable** for the case it was built for. + - **Verified after, on a fresh document with no reload and no second keystroke**: verdict `pair`, selection read aloud in pt-PT, `ver` captured as `lang=pair` with the English sense as its meaning and its Portuguese phonetic, the garden showing `português` on the two pair cards and nothing on the four English ones, and `ver` → pt-PT beside `yesterday` → en-US in the same list. + - **The lesson worth keeping**: every one of 1–4 sat behind tests that passed, because the fixtures were written by the same hand as the code and all argued their own case. The marker list is the sharpest — its test corpus was prose *selected for markers*. + - **Left standing and now visible on a flipped document**: the companion's offline grammar tips are English rules and fire nonsense on Portuguese ("For the past, use past tense: go → went"), the suggestion cards' own chrome ("Accept", "Grammar", "Accept all Grammar (3)") is English while the rest of the UI is bilingual, and **the es Piper container runs on the VPS with no `TTS_ENDPOINT_ES`/`TTS_VOICE_ES` in `.env`** — Spanish read-aloud would 404 to the browser voice. None is Phase 28's, all three are one-liners. - 2026-07-28: **Phase 28 finished — the garden learns which language a card is in, and read-aloud stops guessing** (user: "let's continue the build plan"; step (c), the last of the phase, code only). The two parked questions turned out to have different answers, and the user took the garden one: **tag every card and review all of them**, rather than filter the queue to the half she is learning — the words she met while writing Portuguese are still words she met. Migration `0018_vocab_lang` mirrors `documents.doc_lang` onto `vocab_words`, set server-side from the row-scoped ownership lookup capture was already running, so the tag is free and unforgeable; the one subtlety is that `lang` travels with `doc_id` or not at all, because a lookup from the search box carries no verdict and must not relabel a card. **Read-aloud was the larger surprise: the verdict was server-only.** `detectLang` routed Han/kana to Chinese and *everything else to en-US*, which made the zh pair accidentally right and every Latin pair wrong — a Portuguese selection read by a US English voice. `doc_lang` now rides on the document JSON (read-only, refreshed by the auto-save the editor already makes, so it trails a flip by one save) and `docLang(text, verdict)` answers for a passage lifted out of it. **The script test deliberately still wins over the verdict**: Chinese quoted in an English document is `'en'` and must not be spelled out one "Chinese letter" at a time, while an English sentence inside Portuguese prose is undetectable by construction — which is exactly why the document decides. One thing worth remembering from the build: an extra bind argument passed to `Exec` against a SQL string that never got its matching `?` **did not error** — the planted card came back with an empty `lang` and only the test caught it. Not seen in a browser; not deployed. - 2026-07-28: **Phase 28 finished — the two steps that point the other way** (user: "let's continue the build plan"; the plan's own next items were 28's unbuilt (b), code only). Both steps are about direction, and both turned out to have a wrong answer that looks right. **`isTranslation` could not simply be read backwards**: swapping its two halves would have called every genuine Portuguese correction inside a Portuguese document a translation, because `readsAsEnglish` is a deliberately low bar — Latin letters, not swamped by another script — that Portuguese clears as easily as English does. The flipped direction uses `sentenceLang` from doclang.go instead, where English has its own curated marker list and has to out-evidence the pair language to win; the English-document path is untouched, and the test file now pins both directions with the Portuguese-in-Portuguese case as the one it exists for. **The translate tap-through's whole observable change is a model call that no longer happens**: it now recovers the explanation's language by re-running `targetFor` rather than assuming the pair, which produces exactly today's answer in every case except the one that was broken — the Portuguese writer whose explanation already arrived in Portuguese, previously round-tripped through the model into Portuguese again. It answers `""` there, and the client's existing `res.translation.trim() || explanation` fallback seeds the bubble with the explanation itself, so the fix needed no frontend change at all. It deliberately does *not* render that explanation into English on the grounds that English is technically the other half: an unasked-for rendering into the language she is practising is noise, not a seed, and Ask Petal is where she can ask for it. What is left of the phase is (c) — the garden's language tagging, which touches stored rows, and read-aloud, which has not been traced. Not seen in a browser; not deployed. - 2026-07-28: **Phase 27 — the keystroke that isn't one** (user asked to continue the build plan; the plan's own next item was Phase 26's unbuilt IME guards, named there as the likeliest thing to be wrong the first time anyone types Chinese into Petal for real). **The bug is that a composition is not a keystroke**: the pinyin goes into the document as it is typed, a candidate window sits over it, and all three decoration layers recompute from the live document on every change — rewriting the DOM around the node the browser is composing in, which is what eats half-typed input. **The fix is to hold the redraws, not skip them**: a rebuild that falls due mid-composition marks itself stale and its decorations are *mapped through the transaction*, so they travel with the text growing under them and land correct the moment the composition ends. Skipping would have left every highlight a character behind for as long as she kept typing — the same wrongness, arriving quietly. **The composition flag is read from the state *before* the transaction**, which makes it independent of plugin ordering: the flag was set by compositionstart, not by the transaction being applied, and the end transaction is the one deliberate exception or nothing would ever release. **One piece of timing genuinely matters and is written down where it happens**: a custom `handleDOMEvents` handler runs *before* ProseMirror's own, and ProseMirror's compositionend queues the composition's last DOM changes as a microtask — so the release is a macrotask late, and the held rebuild sees the committed 公园 rather than the *gongyuan* it replaced. If the flush's transaction arrives first it rebuilds anyway; both orders land, which is what makes it not a race. **One thing was checked rather than assumed and turned out already handled**: Tiptap's input-rule plugin returns early while composing — which matters more here than in an English app, because pinyin uses an apostrophe as a syllable separator (xi'an → 西安) and Typography rewrites every `'` into a curly `’`. **The save is deliberately not gated and the analysis is**: a tablet keyboard can hold one composition open for a whole sentence, and Petal never makes writing wait for anything, so `EditorChange` carries the flag, auto-save ignores it, and the checkpoint/rule pack/companion wait — with one more change emitted the instant the composition commits, so nothing is skipped, only deferred by the length of a word. **And four places were quietly stealing keys from the IME**: the Find bar (Enter = next match, Escape = close), the tag picker (Enter = create, in a field whose purpose is a name typed in Chinese), Ask Petal's chat box (the field she types Mandarin into, where Enter sends), and distraction-free mode's global Escape — pressed to fix a wrong candidate, it popped the sidebar back. tsc, vite, go build/vet/test, **vitest 296/296** (12 new, driving real plugin state through a real composition: pinyin in, candidate committed, composition ended). **Then verified in a real browser** (Chrome extension, throwaway DB on :8093): the events are dispatched rather than typed, since this box has no IME, but they drive ProseMirror's own composition machinery, so every layer takes the real path. `gongyuan` typed into the document mid-composition with **the caret moved off it** — so the caret exemption cannot be the explanation — stays un-underlined while `parc` keeps its underline, and underlines on compositionend with no further keystroke. *did a mistake* highlights in 800 ms typed normally and **not at all** after 1.2 s inside a composition, then appears the instant it commits. The save is not held, proven from the other side: with the composition still open and never ended, the server's copy already read `… STILLCOMPOSING`. Escape (`isComposing` and legacy 229) leaves distraction-free mode alone while a real Escape restores the sidebar; the Find bar counter goes 1/4, 1/4, 2/4. **Then with a real IME**, because the user installed one and typed — the only way it could be done: Chrome here runs natively on Wayland, so no key can be injected into its window, and the extension's keys arrive over CDP, which bypasses the OS input method entirely. 很漂亮, composed over **11.6 seconds and 12 candidate changes** (和 → 很 → 狠批 → 很票 → 很漂亮 …), with the preedit genuinely living in the document the whole time — and the underline set never moved, the word committed clean, and English typed afterwards flagged immediately. **The accidental control is the best part**: she typed `henpiaoliang` as plain letters before switching the IME on and it was underlined, so the same syllables appear twice in one sentence, flagged on the path that isn't a composition and untouched on the path that is. ⚠️ **The counterfactual is still unproven** — a clean run with the guard in place can't distinguish "held correctly" from "would have been fine anyway"; settling that means a build with the guard reverted and another minute of her typing. **Not deployed** (no migration — a rebuild, along with Phases 24–26).