diff --git a/docs/branch-review-records/f90327a7dcc408246c5531f97d430de3ada668553127d8034ecfc152d00ad207.record.md b/docs/branch-review-records/f90327a7dcc408246c5531f97d430de3ada668553127d8034ecfc152d00ad207.record.md new file mode 100644 index 0000000000..6368706e48 --- /dev/null +++ b/docs/branch-review-records/f90327a7dcc408246c5531f97d430de3ada668553127d8034ecfc152d00ad207.record.md @@ -0,0 +1 @@ +| 2026-08-17 | claude/ledger-reconcile-issues-c1trbj | 1cff4ddbdf8b700f41ba82c9333034d009ad7e29 | issues:reconcile after #2023/#2024, redone against fresh main | 28 requests applied (22 original + 6 landed on main mid-session: 3 cancel, 2 update, 1 done); redone as one clean transaction after check:ledger-write-discipline rejected the earlier 22-only partial transaction against the advanced base | check:outstanding-issues pass; check:ledger-write-discipline pass against merge-base f5b0932914eb | diff --git a/docs/outstanding-issues-inbox/0649db55-c4ea-4683-b51c-4c6679754f81.json b/docs/outstanding-issues-inbox/applied/0649db55-c4ea-4683-b51c-4c6679754f81.json similarity index 100% rename from docs/outstanding-issues-inbox/0649db55-c4ea-4683-b51c-4c6679754f81.json rename to docs/outstanding-issues-inbox/applied/0649db55-c4ea-4683-b51c-4c6679754f81.json diff --git a/docs/outstanding-issues-inbox/1d35d652-d833-4934-890c-84b6837581cd.json b/docs/outstanding-issues-inbox/applied/1d35d652-d833-4934-890c-84b6837581cd.json similarity index 100% rename from docs/outstanding-issues-inbox/1d35d652-d833-4934-890c-84b6837581cd.json rename to docs/outstanding-issues-inbox/applied/1d35d652-d833-4934-890c-84b6837581cd.json diff --git a/docs/outstanding-issues-inbox/2bfae2cf-d91e-4617-a0cb-bb8adbbad4fc.json b/docs/outstanding-issues-inbox/applied/2bfae2cf-d91e-4617-a0cb-bb8adbbad4fc.json similarity index 100% rename from docs/outstanding-issues-inbox/2bfae2cf-d91e-4617-a0cb-bb8adbbad4fc.json rename to docs/outstanding-issues-inbox/applied/2bfae2cf-d91e-4617-a0cb-bb8adbbad4fc.json diff --git a/docs/outstanding-issues-inbox/2f5fbcde-438b-47ab-a1b7-01309802d935.json b/docs/outstanding-issues-inbox/applied/2f5fbcde-438b-47ab-a1b7-01309802d935.json similarity index 100% rename from docs/outstanding-issues-inbox/2f5fbcde-438b-47ab-a1b7-01309802d935.json rename to docs/outstanding-issues-inbox/applied/2f5fbcde-438b-47ab-a1b7-01309802d935.json diff --git a/docs/outstanding-issues-inbox/314e5e31-7d15-46f2-adb0-c4d113bc589d.json b/docs/outstanding-issues-inbox/applied/314e5e31-7d15-46f2-adb0-c4d113bc589d.json similarity index 100% rename from docs/outstanding-issues-inbox/314e5e31-7d15-46f2-adb0-c4d113bc589d.json rename to docs/outstanding-issues-inbox/applied/314e5e31-7d15-46f2-adb0-c4d113bc589d.json diff --git a/docs/outstanding-issues-inbox/370d002c-5a65-4c99-bf5b-4146c59d5dc8.json b/docs/outstanding-issues-inbox/applied/370d002c-5a65-4c99-bf5b-4146c59d5dc8.json similarity index 100% rename from docs/outstanding-issues-inbox/370d002c-5a65-4c99-bf5b-4146c59d5dc8.json rename to docs/outstanding-issues-inbox/applied/370d002c-5a65-4c99-bf5b-4146c59d5dc8.json diff --git a/docs/outstanding-issues-inbox/38e53f36-b48f-4607-b5a0-56efde6dbb3b.json b/docs/outstanding-issues-inbox/applied/38e53f36-b48f-4607-b5a0-56efde6dbb3b.json similarity index 100% rename from docs/outstanding-issues-inbox/38e53f36-b48f-4607-b5a0-56efde6dbb3b.json rename to docs/outstanding-issues-inbox/applied/38e53f36-b48f-4607-b5a0-56efde6dbb3b.json diff --git a/docs/outstanding-issues-inbox/46750cbf-f4f4-4a00-ab93-bdac644afbde.json b/docs/outstanding-issues-inbox/applied/46750cbf-f4f4-4a00-ab93-bdac644afbde.json similarity index 100% rename from docs/outstanding-issues-inbox/46750cbf-f4f4-4a00-ab93-bdac644afbde.json rename to docs/outstanding-issues-inbox/applied/46750cbf-f4f4-4a00-ab93-bdac644afbde.json diff --git a/docs/outstanding-issues-inbox/52d0dbcc-43e0-4660-89c5-9dbd6e080910.json b/docs/outstanding-issues-inbox/applied/52d0dbcc-43e0-4660-89c5-9dbd6e080910.json similarity index 100% rename from docs/outstanding-issues-inbox/52d0dbcc-43e0-4660-89c5-9dbd6e080910.json rename to docs/outstanding-issues-inbox/applied/52d0dbcc-43e0-4660-89c5-9dbd6e080910.json diff --git a/docs/outstanding-issues-inbox/55cb57e8-7e6e-42f8-bd79-81b81a01ef3a.json b/docs/outstanding-issues-inbox/applied/55cb57e8-7e6e-42f8-bd79-81b81a01ef3a.json similarity index 100% rename from docs/outstanding-issues-inbox/55cb57e8-7e6e-42f8-bd79-81b81a01ef3a.json rename to docs/outstanding-issues-inbox/applied/55cb57e8-7e6e-42f8-bd79-81b81a01ef3a.json diff --git a/docs/outstanding-issues-inbox/55f0d5de-8642-487a-b063-44c1c0bc0921.json b/docs/outstanding-issues-inbox/applied/55f0d5de-8642-487a-b063-44c1c0bc0921.json similarity index 100% rename from docs/outstanding-issues-inbox/55f0d5de-8642-487a-b063-44c1c0bc0921.json rename to docs/outstanding-issues-inbox/applied/55f0d5de-8642-487a-b063-44c1c0bc0921.json diff --git a/docs/outstanding-issues-inbox/64dd6966-0697-4b70-a7ee-c8d526729404.json b/docs/outstanding-issues-inbox/applied/64dd6966-0697-4b70-a7ee-c8d526729404.json similarity index 100% rename from docs/outstanding-issues-inbox/64dd6966-0697-4b70-a7ee-c8d526729404.json rename to docs/outstanding-issues-inbox/applied/64dd6966-0697-4b70-a7ee-c8d526729404.json diff --git a/docs/outstanding-issues-inbox/6602201c-14f2-46e9-9407-36e56efe13d5.json b/docs/outstanding-issues-inbox/applied/6602201c-14f2-46e9-9407-36e56efe13d5.json similarity index 100% rename from docs/outstanding-issues-inbox/6602201c-14f2-46e9-9407-36e56efe13d5.json rename to docs/outstanding-issues-inbox/applied/6602201c-14f2-46e9-9407-36e56efe13d5.json diff --git a/docs/outstanding-issues-inbox/6c0d5004-45a4-42c8-b001-92428f76a0b0.json b/docs/outstanding-issues-inbox/applied/6c0d5004-45a4-42c8-b001-92428f76a0b0.json similarity index 100% rename from docs/outstanding-issues-inbox/6c0d5004-45a4-42c8-b001-92428f76a0b0.json rename to docs/outstanding-issues-inbox/applied/6c0d5004-45a4-42c8-b001-92428f76a0b0.json diff --git a/docs/outstanding-issues-inbox/739887b7-24d6-4ea8-9aad-e7700a537913.json b/docs/outstanding-issues-inbox/applied/739887b7-24d6-4ea8-9aad-e7700a537913.json similarity index 100% rename from docs/outstanding-issues-inbox/739887b7-24d6-4ea8-9aad-e7700a537913.json rename to docs/outstanding-issues-inbox/applied/739887b7-24d6-4ea8-9aad-e7700a537913.json diff --git a/docs/outstanding-issues-inbox/7ad31d10-aeef-4635-863d-d6ce9340d916.json b/docs/outstanding-issues-inbox/applied/7ad31d10-aeef-4635-863d-d6ce9340d916.json similarity index 100% rename from docs/outstanding-issues-inbox/7ad31d10-aeef-4635-863d-d6ce9340d916.json rename to docs/outstanding-issues-inbox/applied/7ad31d10-aeef-4635-863d-d6ce9340d916.json diff --git a/docs/outstanding-issues-inbox/7c602a99-9db9-4a56-9d7b-1692db146b5d.json b/docs/outstanding-issues-inbox/applied/7c602a99-9db9-4a56-9d7b-1692db146b5d.json similarity index 100% rename from docs/outstanding-issues-inbox/7c602a99-9db9-4a56-9d7b-1692db146b5d.json rename to docs/outstanding-issues-inbox/applied/7c602a99-9db9-4a56-9d7b-1692db146b5d.json diff --git a/docs/outstanding-issues-inbox/7e454b66-74ae-4e67-8669-1f1fdc2c018c.json b/docs/outstanding-issues-inbox/applied/7e454b66-74ae-4e67-8669-1f1fdc2c018c.json similarity index 100% rename from docs/outstanding-issues-inbox/7e454b66-74ae-4e67-8669-1f1fdc2c018c.json rename to docs/outstanding-issues-inbox/applied/7e454b66-74ae-4e67-8669-1f1fdc2c018c.json diff --git a/docs/outstanding-issues-inbox/81fdf812-4dd6-4fd6-b2e8-f3b730affd8a.json b/docs/outstanding-issues-inbox/applied/81fdf812-4dd6-4fd6-b2e8-f3b730affd8a.json similarity index 100% rename from docs/outstanding-issues-inbox/81fdf812-4dd6-4fd6-b2e8-f3b730affd8a.json rename to docs/outstanding-issues-inbox/applied/81fdf812-4dd6-4fd6-b2e8-f3b730affd8a.json diff --git a/docs/outstanding-issues-inbox/855d03b3-61a8-4a81-ad31-80d996dd581a.json b/docs/outstanding-issues-inbox/applied/855d03b3-61a8-4a81-ad31-80d996dd581a.json similarity index 100% rename from docs/outstanding-issues-inbox/855d03b3-61a8-4a81-ad31-80d996dd581a.json rename to docs/outstanding-issues-inbox/applied/855d03b3-61a8-4a81-ad31-80d996dd581a.json diff --git a/docs/outstanding-issues-inbox/88628da6-881d-4947-a710-56093c5503fa.json b/docs/outstanding-issues-inbox/applied/88628da6-881d-4947-a710-56093c5503fa.json similarity index 100% rename from docs/outstanding-issues-inbox/88628da6-881d-4947-a710-56093c5503fa.json rename to docs/outstanding-issues-inbox/applied/88628da6-881d-4947-a710-56093c5503fa.json diff --git a/docs/outstanding-issues-inbox/a3ba2812-2984-4fad-b959-6dfe69fb6aa8.json b/docs/outstanding-issues-inbox/applied/a3ba2812-2984-4fad-b959-6dfe69fb6aa8.json similarity index 100% rename from docs/outstanding-issues-inbox/a3ba2812-2984-4fad-b959-6dfe69fb6aa8.json rename to docs/outstanding-issues-inbox/applied/a3ba2812-2984-4fad-b959-6dfe69fb6aa8.json diff --git a/docs/outstanding-issues-inbox/a81b68eb-6b26-46db-81fa-90279a599886.json b/docs/outstanding-issues-inbox/applied/a81b68eb-6b26-46db-81fa-90279a599886.json similarity index 100% rename from docs/outstanding-issues-inbox/a81b68eb-6b26-46db-81fa-90279a599886.json rename to docs/outstanding-issues-inbox/applied/a81b68eb-6b26-46db-81fa-90279a599886.json diff --git a/docs/outstanding-issues-inbox/aa0eb6ae-8f6a-4440-8a15-6ba0d17f812d.json b/docs/outstanding-issues-inbox/applied/aa0eb6ae-8f6a-4440-8a15-6ba0d17f812d.json similarity index 100% rename from docs/outstanding-issues-inbox/aa0eb6ae-8f6a-4440-8a15-6ba0d17f812d.json rename to docs/outstanding-issues-inbox/applied/aa0eb6ae-8f6a-4440-8a15-6ba0d17f812d.json diff --git a/docs/outstanding-issues-inbox/c156f1c0-8d83-460a-b1b5-ba970d664405.json b/docs/outstanding-issues-inbox/applied/c156f1c0-8d83-460a-b1b5-ba970d664405.json similarity index 100% rename from docs/outstanding-issues-inbox/c156f1c0-8d83-460a-b1b5-ba970d664405.json rename to docs/outstanding-issues-inbox/applied/c156f1c0-8d83-460a-b1b5-ba970d664405.json diff --git a/docs/outstanding-issues-inbox/db498cc1-c516-4141-9837-15fc9ef30684.json b/docs/outstanding-issues-inbox/applied/db498cc1-c516-4141-9837-15fc9ef30684.json similarity index 100% rename from docs/outstanding-issues-inbox/db498cc1-c516-4141-9837-15fc9ef30684.json rename to docs/outstanding-issues-inbox/applied/db498cc1-c516-4141-9837-15fc9ef30684.json diff --git a/docs/outstanding-issues-inbox/e54f04f8-09bb-4887-b3e4-a5cd99958798.json b/docs/outstanding-issues-inbox/applied/e54f04f8-09bb-4887-b3e4-a5cd99958798.json similarity index 100% rename from docs/outstanding-issues-inbox/e54f04f8-09bb-4887-b3e4-a5cd99958798.json rename to docs/outstanding-issues-inbox/applied/e54f04f8-09bb-4887-b3e4-a5cd99958798.json diff --git a/docs/outstanding-issues-inbox/ff73ce15-6890-44a1-8ddb-2919306a0e86.json b/docs/outstanding-issues-inbox/applied/ff73ce15-6890-44a1-8ddb-2919306a0e86.json similarity index 100% rename from docs/outstanding-issues-inbox/ff73ce15-6890-44a1-8ddb-2919306a0e86.json rename to docs/outstanding-issues-inbox/applied/ff73ce15-6890-44a1-8ddb-2919306a0e86.json diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index 3b89830149..f63a0b2c52 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -77,33 +77,28 @@ removed after current-main verification; it is not missing recommended work. | 22 | `#098` | A3 | High — test infrastructure | Before `#099` or `#101`; it is their enabler | 2–4 hours | Generalise the answer-route preamble guard into a counting-proxy round-trip budget harness over the existing offline fixtures. Must enforce admission-before-scope, never the reverse. No providers, no DB. Stop if it would require live credentials. | | 23 | `#102` | A3 | Operator — Supabase + Specialist | Next approved index window, after the ordering question is settled | 1–2 hours plus apply | Author the migration (operator SQL alone never reaches staging/DR/local replay), then apply → mirror `schema.sql` → regenerate drift manifest → register `required_indexes`. **Stop:** the RAG-path index is canary-gated, and ordering `fetchDocumentTitleAliasRows`'s unordered `.limit(12)` does not lift that — an imposed order can select a different twelve, so it is a second canary-gated change, not a way out of the first. The byte-identical claim was retracted. | | 24 | `#099` | A3 | Specialist — answer path | After `#098` | Half a day per sub-item | Remaining fixed per-request round trips: the 8 `setCachedSearch` deferrals (abort semantics + mutation window), the anonymous subject+global limiter pair (needs a new atomic RPC first), and proxy→route identity duplication. Stop before hand-authoring locking SQL. | -| 25 | `#162` | A3 | High — frontend/UI | When starting the mode search redesign package | 0.5–1.5 days | Redesign `/tools?q=` as Compact Results Instrument (direction A): query-as-H1, one composer, dense tool rows, demote cross-mode cards, remove success-green filter banner and home hero on results. Comps in `public/mockups/mode-page-redesign-2026-07/tools-search/`. Verify phone+desktop chrome ownership and `verify:phone-chrome` / focused UI. Stop if scope expands into Tools home redesign without an explicit ask. | -| 26 | `#090` | A3 | High — eslint toolchain | When ESLint 10 plugin peers are compatible | blocked; revisit monthly | Upgrade the eslint ecosystem to clear remaining dev-scoped high advisories — full `npm audit` reports zero high advisories from the eslint toolchain. | -| 27 | `#100` | A3 | Specialist — answer streaming | After offline Phase 0/1 design proof | provider-gated rollout | Buffered answer generation has no incremental verified delivery — [`verified-answer-incremental-delivery-design.md`](verified-answer-incremental-delivery-design.md) records the clinical-governance decision and staged co… | -| 28 | `#150` | Optional | Operator — review tooling | Next CodeRabbit billing/policy decision | 30–60 min decision | CodeRabbit reviewed none of a full day's PRs; spending cap reached — the repo's second automated reviewer is either funded or acknowledged as absent, rather than appearing to review while skipping. | -| 29 | `#165` | A2 | High — clinical UI | Next answer-home UX pass | 0.5–1 day | Adopt a consolidated answer-home notice block — the studies exist, nothing adopts them — the answer hero states its safety obligation, its scope, and its verification requirement as one block in one voice. | -| 30 | `#168` | A3 | High — ledger architecture | With #156 / id-scheme redesign | design first | Sequential issue ids force every concurrent append to conflict — two sessions can append to this ledger at the same time without conflicting. | -| 31 | `#169` | A3 | High — git hygiene | Next branch cleanup batch | 1–2 hours | Local branches carry work that exists on no remote — committed work is not lost when a machine or worktree is reclaimed. | -| 32 | `#175` | A2 | Operator — clinical data + Standard | Next therapy catalogue curation window | 2–4 hours | Therapy modality is now null on all 205 records and needs curation or removal — the Therapy detail and recommend screens either show a curated modality or stop carrying the field at all. | -| 33 | `#036` | Optional | Specialist — privacy/schema | When visibility model is redesigned | design + migration | No explicit `is_public` visibility flag on documents — Public-corpus visibility is implicit: `owner_id IS NULL` on an `indexed` document (`resolveSearchScope`). The `metadata.public_corpus` marker is written by the prom… | -| 34 | `#101` | A3 | Specialist — RAG/retrieval | After #098 harness + canary approval | canary-gated | Canary-gated retrieval parallelisation candidates — metadata and memory hydration shipped in PR #1474; visual hydration, scope enumeration, typeahead caching, and universal-search coalescing remain, each behind the RAG flag and live-canary criteria. | -| 35 | `#190` | A3 | Specialist — RAG structure | On explicit X3 go-ahead | 1 PR per extraction unit | X3: Finish rag.ts monolith decomposition — `src/lib/rag/rag.ts` is decomposed into focused modules per `docs/maturity-backlog-workorders.md` X3, with existing offline RAG contracts green. | -| 36 | `#191` | A3 | Operator — DB + Specialist | Approved live-DB window only | provider-gated | X5: ACL-migration consolidation (provider-gated) — ACL-related migrations are consolidated per maturity work-order X5 without weakening owner-scope/RLS. | -| 37 | `#192` | A3 | High — test coverage | Next coverage-floor pass | 0.5–1 day | X6: Raise clinical/retrieval/answer coverage floors — coverage floors for clinical, retrieval, and answer domains meet the maturity X6 targets with CI enforcing them. | -| 38 | `#193` | A3 | High — src/lib structure | After/with X3 non-protected clusters | 1 PR per cluster | X7: Complete the remaining src/lib domain-directory reorg — remaining `src/lib` clusters sit in their domain directories per X7 follow-on to X2. | -| 39 | `#195` | A3 | Operator — GitHub maintainer | Maintainer UI window | 30–60 min | M1: Repo-host hardening (branch protection and required checks) — GitHub branch-protection rulesets and required checks match audit §8 / maturity M1. | -| 40 | `#183` | A2 | Operator — Sentry + Specialist | Next approved observability window with SENTRY_AUTH_TOKEN | 1–2 hours | Create Sentry metric alert for production DB span p95 > 500ms (`span.op:db`, environment production). **Stop:** no secret printing; blocked until token/env available. | -| 41 | `#206` | A2 | Specialist — answer UI contract | With AnswerState producer work (`#207`) | 2–4 hours | `partial_retrieval` has no app-facing producer — decide RAG contract vs UI-only mapping before AnswerCard. **Stop:** no retrieval behaviour change without RAG flag. | -| 42 | `#211` | A3 | High — TypeScript strictness | Dedicated migration branch | multi-PR | Plan and start `noUncheckedIndexedAccess` migration (1266 errors); highest-risk files first. **Stop:** do not flip the flag on main without a staged plan. | -| 43 | `#212` | A3 | High — runtime validation | After highest-risk cast inventory | multi-PR | Replace `as unknown as` and unvalidated `JSON.parse` with Zod/guards at trust boundaries. **Stop:** RAG/provider boundaries need clinical/privacy care. | -| 44 | `#222` | A3 | High — headers / search chrome | During headers redesign decision | 2–4 hours | Decide whether mode-home-template / search-results-header-band are in PageHeader scope or permanently out. **Stop:** do not flatten phone composer ownership. | -| 45 | `#235` | A3 | High — design-system evidence | Next warmed local proof-shot pass | 1–2 hours | Capture missing ADOPTION.md §7 proof shots for adopted surfaces. **Stop:** not visual-baseline PNGs (`#118`); no Playwright snapshot commit. | -| 46 | `#237` | A3 | High — design-system a11y | Before freezing Linux visual baselines (#242) | 30–60 min | Eyeball low-confidence AccessibleTable densities at 320px; MissingValue phrases must remain readable. **Gate:** visual spot-check only. **Stop:** do not abbreviate MissingValue to a dash. | -| 47 | `#238` | A3 | High — overlays/UI | After Sheet portal default change (#1616) | 30–60 min | Visual pass for Sheet portal default on settings, sidebar, and answer overlays under OverlayRoot. **Stop:** do not revert portal default without evidence. | -| 48 | `#239` | Optional | High — phone chrome | When phone orientation QA is available | 15–30 min | Manual phone rotation check for ResizeObserver-only phone chrome reserve. **Gate:** `verify:phone-chrome` still owns automated coverage. **Stop:** do not widen reserve heuristics without reproduction. | -| 49 | `#240` | Optional | High — design tokens | Next design-owner review | 15–30 min | Confirm tooltip visual hard-clip asymmetry with design owner (sr-only keeps full text). **Stop:** no product change without that confirmation. | -| 50 | `#242` | A2 | High — design-system baselines | After human review of Linux baselines | 1–2 hours | Commit approved Linux visual baselines and promote adoption not-committed → committed. **Stop:** never commit baselines from an unreviewed machine run. | -| 51 | `#248` | A2 | Operator — Supabase + Specialist | After PR #1614 symptom repair; approved live/history window | 1–2 hours | Investigate why 20260705180000 search-health indexes were missing on live despite applied history; decide if drift checks should catch this class. **Stop:** no hosted mutation without approval. | +| 25 | `#090` | A3 | High — eslint toolchain | When ESLint 10 plugin peers are compatible | blocked; revisit monthly | Upgrade the eslint ecosystem to clear remaining dev-scoped high advisories — full `npm audit` reports zero high advisories from the eslint toolchain. | +| 26 | `#100` | A3 | Specialist — answer streaming | After offline Phase 0/1 design proof | provider-gated rollout | Buffered answer generation has no incremental verified delivery — [`verified-answer-incremental-delivery-design.md`](verified-answer-incremental-delivery-design.md) records the clinical-governance decision and staged co… | +| 27 | `#150` | Optional | Operator — review tooling | Next CodeRabbit billing/policy decision | 30–60 min decision | CodeRabbit reviewed none of a full day's PRs; spending cap reached — the repo's second automated reviewer is either funded or acknowledged as absent, rather than appearing to review while skipping. | +| 28 | `#165` | A2 | High — clinical UI | Next answer-home UX pass | 0.5–1 day | Adopt a consolidated answer-home notice block — the studies exist, nothing adopts them — the answer hero states its safety obligation, its scope, and its verification requirement as one block in one voice. | +| 29 | `#168` | A3 | High — ledger architecture | With #156 / id-scheme redesign | design first | Sequential issue ids force every concurrent append to conflict — two sessions can append to this ledger at the same time without conflicting. | +| 30 | `#169` | A3 | High — git hygiene | Next branch cleanup batch | 1–2 hours | Local branches carry work that exists on no remote — committed work is not lost when a machine or worktree is reclaimed. | +| 31 | `#175` | A2 | Operator — clinical data + Standard | Next therapy catalogue curation window | 2–4 hours | Therapy modality is now null on all 205 records and needs curation or removal — the Therapy detail and recommend screens either show a curated modality or stop carrying the field at all. | +| 32 | `#036` | Optional | Specialist — privacy/schema | When visibility model is redesigned | design + migration | No explicit `is_public` visibility flag on documents — Public-corpus visibility is implicit: `owner_id IS NULL` on an `indexed` document (`resolveSearchScope`). The `metadata.public_corpus` marker is written by the prom… | +| 33 | `#101` | A3 | Specialist — RAG/retrieval | After #098 harness + canary approval | canary-gated | Canary-gated retrieval parallelisation candidates — metadata and memory hydration shipped in PR #1474; visual hydration, scope enumeration, typeahead caching, and universal-search coalescing remain, each behind the RAG flag and live-canary criteria. | +| 34 | `#190` | A3 | Specialist — RAG structure | On explicit X3 go-ahead | 1 PR per extraction unit | X3: Finish rag.ts monolith decomposition — `src/lib/rag/rag.ts` is decomposed into focused modules per `docs/maturity-backlog-workorders.md` X3, with existing offline RAG contracts green. | +| 35 | `#191` | A3 | Operator — DB + Specialist | Approved live-DB window only | provider-gated | X5: ACL-migration consolidation (provider-gated) — ACL-related migrations are consolidated per maturity work-order X5 without weakening owner-scope/RLS. | +| 36 | `#193` | A3 | High — src/lib structure | After/with X3 non-protected clusters | 1 PR per cluster | X7: Complete the remaining src/lib domain-directory reorg — remaining `src/lib` clusters sit in their domain directories per X7 follow-on to X2. | +| 37 | `#195` | A3 | Operator — GitHub maintainer | Maintainer UI window | 30–60 min | M1: Repo-host hardening (branch protection and required checks) — GitHub branch-protection rulesets and required checks match audit §8 / maturity M1. | +| 38 | `#183` | A2 | Operator — Sentry + Specialist | Next approved observability window with SENTRY_AUTH_TOKEN | 1–2 hours | Create Sentry metric alert for production DB span p95 > 500ms (`span.op:db`, environment production). **Stop:** no secret printing; blocked until token/env available. | +| 39 | `#206` | A2 | Specialist — answer UI contract | With AnswerState producer work (`#207`) | 2–4 hours | `partial_retrieval` has no app-facing producer — decide RAG contract vs UI-only mapping before AnswerCard. **Stop:** no retrieval behaviour change without RAG flag. | +| 40 | `#211` | A3 | High — TypeScript strictness | Dedicated migration branch | multi-PR | Plan and start `noUncheckedIndexedAccess` migration (1266 errors); highest-risk files first. **Stop:** do not flip the flag on main without a staged plan. | +| 41 | `#222` | A3 | High — headers / search chrome | During headers redesign decision | 2–4 hours | Decide whether mode-home-template / search-results-header-band are in PageHeader scope or permanently out. **Stop:** do not flatten phone composer ownership. | +| 42 | `#235` | A3 | High — design-system evidence | Next warmed local proof-shot pass | 1–2 hours | Capture missing ADOPTION.md §7 proof shots for adopted surfaces. **Stop:** not visual-baseline PNGs (`#118`); no Playwright snapshot commit. | +| 43 | `#239` | Optional | High — phone chrome | When phone orientation QA is available | 15–30 min | Manual phone rotation check for ResizeObserver-only phone chrome reserve. **Gate:** `verify:phone-chrome` still owns automated coverage. **Stop:** do not widen reserve heuristics without reproduction. | +| 44 | `#240` | Optional | High — design tokens | Next design-owner review | 15–30 min | Confirm tooltip visual hard-clip asymmetry with design owner (sr-only keeps full text). **Stop:** no product change without that confirmation. | +| 45 | `#242` | A2 | High — design-system baselines | After human review of Linux baselines | 1–2 hours | Commit approved Linux visual baselines and promote adoption not-committed → committed. **Stop:** never commit baselines from an unreviewed machine run. | +| 46 | `#248` | A2 | Operator — Supabase + Specialist | After PR #1614 symptom repair; approved live/history window | 1–2 hours | Investigate why 20260705180000 search-health indexes were missing on live despite applied history; decide if drift checks should catch this class. **Stop:** no hosted mutation without approval. | @@ -128,7 +123,7 @@ removed after current-main verification; it is not missing recommended work. | #001 | P2 | task | Semantic reranking still gated off | `RAG_SEMANTIC_RERANK_ENABLED=false` from PR #901. Do not enable until the provider-backed 36/36 retrieval-quality gate **and** an ambiguity-focused canary are explicitly approved and recorded. | `docs/process-hardening.md` (Semantic reranking rollout debt); PR #901 | 2026-07-21 | | #053 | P1 | task | Execute cross-border privacy/legal package | Execute OpenAI and Railway DPAs; decide ZDR and Australian data residency; obtain prompt-cache behavior in writing; review subprocessors; obtain APP 8 and APP 5/1 counsel sign-off. Do not represent the release as privacy-approved or alter final public privacy wording before sign-off. | `docs/openai-cross-border-basis.md`; `docs/privacy-impact-assessment.md` | 2026-07-24 | | #055 | P2 | task | Run one exact-SHA full release and PR gate | Before the next full-confidence release/handoff, record the candidate/PR SHA and run the local/provider release gates, Firefox/WebKit, required hosted CI, and actionable GitHub review-thread closure once. Stop at the first actionable failure and rerun only the repaired smallest gate. | `docs/launch-operator-runbook.md`; `docs/codex-review-protocol.md` | 2026-07-24 | -| #056 | P2 | task | Reconcile the existing staging migration history | `Clinical KB Staging` already exists as a healthy, empty Supabase/Railway tier with distinct secrets and no production clinical data, but it is 24 repository migrations behind (ten earlier history holes plus fourteen after `20260719055623`). In the next approved staging schema window, apply the exact missing migration chain, then re-run indexing, health, identity and data-boundary proof. Do not recreate the environment or copy production clinical documents. | current-main staging verification; `docs/staging-setup.md`; `docs/operator-backlog.md` | 2026-07-27 | +| #056 | P2 | task | Reconcile the existing staging migration history | `Clinical KB Staging` already exists as a healthy, empty Supabase/Railway tier with distinct secrets and no production clinical data, but it is behind the repository migration chain. In the next approved staging schema window, apply the exact missing migration chain, then re-run indexing, health, identity and data-boundary proof. Do not recreate the environment or copy production clinical documents. COUNTS REMEASURED 2026-08-15 (read-only connector query against ikoiolksxqxfxgiyqpnu): the gap is now 26, not 24, and it is widening rather than static. Staging holds 166 rows in supabase_migrations.schema_migrations with latest version 20260719055623; the repository holds 192 migration files, of which 16 - not fourteen - are dated after 20260719055623. So the composition is ten earlier history holes plus sixteen later migrations. The two-migration increase is simply new work landing on main while staging stands still, which is the expected behaviour of an unattended staging tier and not a new defect - but it does mean the apply chain grows every week this row stays open, and any runbook written against the earlier figures is already stale. Re-measure the counts at the start of the window rather than trusting these; they were accurate on 2026-08-15 and will not stay that way. | current-main staging verification; `docs/staging-setup.md`; `docs/operator-backlog.md` | 2026-07-27 | | #057 | P2 | task | Complete staging soak and rollback rehearsal | After #056, run the documented soak and rollback against an exact candidate; retain latency/error/rollback evidence. Stop on unsafe data, identity mismatch, or an unowned rollback decision. | `docs/launch-operator-runbook.md`; `docs/audit/capacity-review.md` | 2026-07-24 | | #011 | P3 | task | Auth DB-connection allocation is operator-only | Supabase Auth (GoTrue) is capped at ~10 absolute DB connections (Supabase perf advisor). Switch to **percentage-based** allocation in the Supabase **dashboard** before the first compute scale-up — **not settable via SQL/MCP** (operator-owned). Verify via a staging soak + an approval-gated read-only advisor re-check. | `docs/auth-connection-cap-runbook.md`; `docs/process-hardening.md` (Known follow-up debts) | 2026-07-21 | | #013 | P3 | rec | Route catalogue weight remains measurement-gated; public field INP is unavailable | Keep the payload work open, but correct the measurement state: on 2026-08-13 the official Chrome UX Report current-record API returned 404 no-data for the psychiatry.tools origin and every reviewed URL (root, Therapy, Documents search, DSM, Forms, Services, Specifiers, and Formulation). Field INP is therefore unavailable because the site does not meet CrUX eligibility/coverage, not unverified and not a pass. The production app deliberately has no browser Sentry bundle, so adding RUM would expand the approved telemetry and privacy envelope. Next: continue lab LCP/TBT and interaction traces; request a separate privacy/operator decision before adding browser RUM, or recheck CrUX after traffic eligibility changes. Do not block route payload fixes waiting for a field dataset that does not exist, and do not infer an INP pass from absence. | session 2026-08-13 official CrUX current-record queries; https://cruxvis.withgoogle.com/ | 2026-07-21 | @@ -145,7 +140,7 @@ removed after current-main verification; it is not missing recommended work. | #039 | P3 | rec | Consolidate catalogue toolbar patterns | **DEPRIORITISED 2026-08-12 (yield review against current main).** States an intent (converge repeated toolbar behaviour) with no measured defect and no named surfaces. Needs a concrete inventory before it is work. Catalogue/search pages have independently evolved filter, sort, result-count and mobile toolbar behavior. Inventory the existing implementations and converge only the repeated interaction contract; do not flatten mode-specific search semantics. | design audit reconciliation; session 2026-07-22 | 2026-07-22 | | #079 | P3 | task | Disposition retained worktrees in bounded cleanup batches | **Outcome:** the retained reconciliation tail is gradually classified without another disruptive all-worktree sweep. **Next:** after the primary checkout is clean and `npm run check:primary-checkout-lease` allows writes, revalidate and remove the twenty clean redundant candidates recorded on 2026-07-30 with `branch-cleanup-deletion-pending`; then process no more than ten further worktrees per explicitly scheduled pass using current owner/process metadata, open-PR state, exact review-ledger coverage, ancestry, and cherry-pick-aware content proof. **Success:** remove only clean, inactive, bundled worktrees whose content is merged or explicitly rejected; record every disposition and retain recovery evidence. **Stop:** preserve dirty, active, secret-bearing, post-freeze, paused, or ambiguous work and never use reset, force deletion, broad clean, or process killing. | final reconciliation inventory retained 104 independent worktrees; session 2026-07-24; 2026-07-30 bounded review found 20 redundant candidates across two bounded batches but the primary-dirty write lease blocked removal | 2026-07-30 | | #090 | P3 | task | Upgrade the eslint ecosystem to clear remaining dev-scoped high advisories | **DEPRIORITISED 2026-08-12 (yield review against current main).** Blocked upstream on three plugins publishing ESLint 10 support, and the advisories are dev-scoped only with no runtime or user impact. A recheck reminder, not work. **Outcome:** full `npm audit` reports zero high advisories from the eslint toolchain. **Blocked 2026-07-30:** the stable ecosystem still has no compatible ESLint 10 set. `eslint-config-next@16.2.12` permits ESLint 10 but bundles `eslint-plugin-react@7.37.5`, `eslint-plugin-import@2.32.0`, and `eslint-plugin-jsx-a11y@6.10.2`; each plugin's published peer range still ends at ESLint 9, and the React plugin retains the previously reproduced removed-context-API crash. Keep the Dependabot major hold and ESLint `9.39.5`; do not force an invalid peer graph merely to make the audit report green. **Next:** recheck after those three plugins publish stable ESLint 10 support, then upgrade eslint and the complete plugin/config set together. Residual highs (`@eslint/config-array`, `@eslint/eslintrc`, `eslint`, `eslint-config-next`, `eslint-plugin-import`, `eslint-plugin-jsx-a11y`, plus the advisory's numeric `<=5.0.7` hit on the unused `brace-expansion@1.1.16` / `2.1.2` maintenance lines that still ship an unpatched `main`) cascade from this dev-only toolchain. **Success:** peer-valid install, `npm run lint` clean, `verify:cheap` green, full-audit highs cleared, no rule-config regressions. **Stop:** do not use `npm audit fix --force` or override plugin peer ranges. Production `npm audit --omit=dev` is already clean after the exceljs `archiver@8` / `unzipper@0.12.5` overrides on PR #1314. | stable npm metadata recheck 2026-07-30; session 2026-07-28 brace-expansion triage (PR #1314) | 2026-07-30 | -| #098 | P2 | task | Offline round-trip budget harness for the hot routes | **Outcome:** per-scenario Supabase round-trip counts are pinned by a test, so an extra round trip on a hot path is a red gate rather than an inference. **Done 2026-07-29:** the measurement gap is closed — `Server-Timing` now covers `auth`/`ratelimit`/`scope` on `/api/answer`, `auth`/`ratelimit`/`search`/`total` on `/api/search`, and `auth`/`ratelimit` on `/api/answer/stream` (previously the route the UI actually calls emitted no header at all). Headers flush before the first SSE frame, so in-stream stages cannot reach a header and must NOT be routed through the governed `progress`/`final` contract. `tests/answer-route-preamble.test.ts` pins admission-before-scope (no scope call while the limiter is pending or after a deny) and the client-disconnect abort signal. **Done 2026-07-30 (PR #1450, `1bff4c78`):** the counting proxy exists and the answer path is budgeted. `tests/helpers/supabase-round-trip-counter.ts` counts on **execution, not construction** — a builder that is never awaited costs zero, one awaited twice costs two — which is the distinction that makes the count mean "requests issued". `tests/rag-round-trip-budget.test.ts` pins two offline answer-path scenarios (a single-source source-only answer, and that trips do not scale with the number of retrieved sources) plus three self-tests of the counter, and is registered in `scripts/fixtures/rag-offline-contract-tests.json` so it runs inside the offline contract rather than only on demand. Verified locally, provider-free: `Test Files 1 passed (1)`, `Tests 5 passed (5)`. Its documented blind spot is worth repeating before anyone cites a budget as total cost: it sees only traffic through the wrapped client, so a trip issued via another client instance, a direct `fetch`, or a provider SDK is invisible to it. **Done 2026-07-30 (search *retrieval core*, not the endpoint):** `tests/search-round-trip-budget.test.ts` pins `searchChunksWithTelemetry` — what `/api/search` calls to retrieve — registered in both the contract fixture and `scripts/rag-offline-contract.mjs`. **Corrected after Codex review on PR #1464:** an earlier version of this row and the test itself claimed to pin `/api/search`. They do not. The route's auth, rate limiting, scope resolution, related-document enrichment and telemetry write are all invisible to this suite, so a round trip added to any of them leaves it green — and the refusal budget below is about *retrieval*, not about an adversarial HTTP request, which still pays the route preamble. **The measured shape is itself the finding:** one search costs **11 round trips** — `rag_aliases` 1, `match_document_chunks_text_v2` **3**, `match_document_table_facts_text_v2` **3**, `get_related_document_metadata_v2` 1, `document_index_quality` 1, `document_images` 2 — so the two text RPCs are each issued three times per search. Pinned by total *and* breakdown, because a refactor swapping one probe for an unrelated query would keep the total at 11 while changing the traffic. Deterministic across three consecutive runs. The refusal budget asserts **zero** Supabase traffic, matching `rag.ts`'s claim that prompt-injection intent is refused before any query issues, and was proven against the broken shape: with a non-refused query it fails on the round-trip assertion (`expected 11 to be +0`), which is why that assertion is ordered ahead of the results assertion. **Note 2026-07-30 (corrected):** a work-branch experiment on `origin/work` (`1f52c704`, reverted in `a0cd00ba`) collapsed both text surfaces offline (budget 11→7) but never reached `main` and was never canaried on those SHAs. Live canaries `30579804611` and `30580564419` ran on unrelated `main` docs commits (`fde68ed4` / `4312a214`) and must not be cited as probe-collapse evidence. Next (b) remains open and still needs a real canary pair before any collapse. **Next:** (a) add the route-level budget this suite does not provide — drive `POST` from `src/app/api/search/route.ts` with counted clients, following the `tests/answer-route-preamble.test.ts` pattern, so a round trip added to the route preamble or post-processing is a red gate; (b) do **not** wholesale-collapse the ×3 sibling lexical variants as the next step — the offline collapse on `origin/work` never landed on `main` with a canary on those SHAs (see Note above), so that avenue is not an approved follow-up; any later latency work must use a materially different approach (for example overlap/parallelism that preserves all three variants) or a real canary pair on the changed tip under the usual RAG gate. (c) `scripts/eval-rag-offline.mjs` and `scripts/test-rag-offline.mjs` remain unwired; the offline contract runner is now the de-facto single home for budgets, so either adopt that explicitly here or wire them. | `docs/audit/latency-audit-2026-07-28.md` measurement plan; `src/lib/server-timing.ts`; `src/lib/answer-stream-contract.ts:18-21` | 2026-07-29 | +| #098 | P2 | task | Offline round-trip budget harness for the hot routes | UPDATE 2026-08-16: residual (a) is complete on current main via commit 563ce4195512b9e623df6ba98762f4a3dc1b9e8e. tests/search-route-round-trip-budget.test.ts now drives POST /api/search with a counted Supabase client and pins the successful-request route budget plus the limiter-only denial path, so route preamble and post-processing traffic are covered. Keep this row open: residual (c) is still unresolved. scripts/eval-rag-offline.mjs and scripts/test-rag-offline.mjs remain unwired, and the new route-level test is not registered in scripts/fixtures/rag-offline-contract-tests.json. Decide explicitly whether the offline contract runner is the canonical home and register this route-level budget there, or wire the legacy scripts. The prior stop against wholesale collapsing the three lexical variants remains unchanged; any later latency change still needs a materially different approach or a real canary pair. Stop: do not close #098 until the runner decision is implemented and recorded. | Commit 563ce4195512b9e623df6ba98762f4a3dc1b9e8e; tests/search-route-round-trip-budget.test.ts; scripts/fixtures/rag-offline-contract-tests.json | 2026-07-29 | | #099 | P2 | task | Remove the remaining fixed per-request round trips | **Outcome:** the answer path stops paying avoidable per-request Supabase round trips. **Done 2026-07-29:** shared-cache-hit promotion deferred off the response path with its mid-request staleness guard intact and documented (`rag.ts:3234`, `rag-cache.ts`); scope resolution overlapped with the rate-limit RPC, signal threaded so a client disconnect finally cancels its paginated queries (`answer/route.ts`). **REFUTED on PR #1377 review — do not retry:** the same pass also overlapped scope with the rate-limit RPC and aborted it on deny, claiming the limiter could "deny for free". It cannot. With caller-supplied `filters` or explicit ids, scope passes its zero-query early returns (`search-scope.ts:242,253`) into the paginated `documents` loop at `:269`, and an `AbortSignal` cancels the client request without un-executing a statement Postgres already began — so throttled traffic kept burning database capacity while collecting 429s, against `capacity-review.md:106-113`'s first-soft-failure warning. Scope is behind admission again, pinned by `tests/answer-route-preamble.test.ts`. Re-attempting the overlap requires a non-database admission gate ahead of the durable limiter first. **Remaining:** (a) the 8 `setCachedSearch` awaits — deferring changes `throwIfAborted` semantics and widens a real mutation window because the clone happens after an `await`, so each branch needs discharging individually; (b) batch the anonymous subject+global rate-limit pair, which needs a NEW atomic RPC modelled on `consume_summary_rate_limits_atomic` and cannot be called until the operator applies it — `Promise.all` is the WRONG fix because it consumes the global bucket even when the subject bucket already denied; (c) stop the proxy and route handler resolving identity twice per authenticated request — no in-process memo can do this (different `Request` objects), so the proxy must forward unspoofable verified claims via a header it controls. Cross-references #011: halving auth resolutions eases the ~10-connection Auth cap that `capacity-review.md:106-113` calls the first hard failure. | `docs/audit/latency-audit-2026-07-28.md` L1-1/L1-3/L1-4; `src/lib/api-rate-limit.ts:276-282`; `src/proxy.ts:125` | 2026-07-29 | | #100 | P2 | rec | Buffered answer generation has no incremental verified delivery | UPDATE 2026-08-13 (PR #1909): Phase 0 offline contract proof and flag-gated Phase 1 server emission implemented (RAG_INCREMENTAL_EVIDENCE_PREVIEW, default false). Remaining: client parsing/rendering phase behind its own flag + verify:ui, then the design's provider-backed acceptance gates before production enablement; Phase 2 stays provider-gated. **Design complete; runtime work remains provider-gated.** [`verified-answer-incremental-delivery-design.md`](verified-answer-incremental-delivery-design.md) records the clinical-governance decision and staged contract: keep the `progress`/`final`/`error` allowlist; disclose bounded, owner-scoped evidence only after the canonical danger-level source-governance refusal permits it, then emit complete answer sections only after each reuses the full production verification boundary; reconcile every preview byte-for-byte with the authoritative `final`; discard all previews on error/cancel/retry; deploy behind separate parse/emission/render flags. Phase 0 contract proof and Phase 1 evidence preview can be developed offline, but visible rollout still needs clinical/browser proof. Phase 2 changes generation architecture and requires explicit approval for answer-quality evals plus a baseline/post live canary pair. **Naive token streaming remains REFUTED:** never re-land `token`, `revising`, provisional prose, or a weaker stream-only verifier. Cross-references #021. | `docs/verified-answer-incremental-delivery-design.md`; `docs/audit/latency-audit-2026-07-28.md` L0-1; `src/lib/answer-stream-contract.ts:18-21` | 2026-07-30 | | #101 | P3 | rec | Canary-gated retrieval parallelisation candidates | **DEPRIORITISED 2026-08-12 (yield review against current main).** Every candidate needs the #098 harness, the RAG flag, live-canary approval, 36/36 retrieval and recall 1.0 — provider spend and clinical-surface risk for latency nobody is currently waiting on. **Outcome:** remaining retrieval parallelisation candidates are explicit after PR #1474 shipped the metadata and memory parallelisation. **Remaining:** visual hydration triples (each migrated path still calls `attachPageVisualEvidence` after hydration — not yet parallelised, see `rag.ts:1442,1811,1857,1959,2194,2281`); nested `await`-in-loop scope enumeration (`search-scope.ts:202,328`); uncached typeahead results (`rag.ts:2698-2711`); and universal-search coalescing (`/api/search` has it; `/api/search/universal` does not). Each changes candidate assembly, truncation, or what the next keystroke returns, so each requires the #098 harness, the RAG flag, explicit live-canary approval, 36/36 retrieval, recall 1.0, and zero per-case reciprocal-rank regressions. Distinct from #001 (semantic rerank). **Completed:** PR #1474 parallelised metadata and memory (`hydrateCandidatesWithMetadataAndMemory`); do not propose that specific change again. **Stop:** no remaining candidate proceeds without its canary gate. | `docs/audit/latency-audit-2026-07-28.md` L2-1/L2-2/L2-8/L1-5; PR #1474 | 2026-07-29 | @@ -153,7 +148,6 @@ removed after current-main verification; it is not missing recommended work. | #117 | P2 | rec | All live mobile routes breach LCP; shared render-blocking CSS and font are the current bottleneck | **Outcome:** `/therapy-compass` mobile LCP lands near the other mobile routes instead of double them. **Measured 2026-07-30** by the new pre-merge Lighthouse budget: mobile LCP 5229 ms, TBT 612 ms, CLS 0.142, against 2123-2460 ms on every other mobile route and 826 ms on desktop — so it is client-side work under mobile CPU/network throttling, not server latency. **Cause before this PR:** `useTherapyData` fetched `/therapy-compass-data/therapies-index.json` (the stable public alias served by a Next rewrite to the thin browse index; 205 records) for the home/search/pathways screens, so the download plus JSON parse sat on the critical path before content painted. **Current split:** home now fetches `public/therapy-compass-data/therapies-home.211dab554c4ec62d.json` (136,288 bytes raw), pathways use the thin browse index, and search loads the full prose corpus (#1471). 90% of the index weight is long-form clinical prose — indications 159 KB (26%), contraindicationsOrCautions 139 KB (23%), bestUsedFor 73 KB (12%), clinicalSummary 67 KB (11%), patientPopulation 59 KB (10%), targetSymptoms 48 KB (8%) — while name, slug, category, tags and setting together are 54 KB (7%). **Remaining decision for search/pathways: rendered on the card, matched by search, or neither.** `therapy-card.tsx` references five of those prose fields and the same index feeds the search screen, so stripping fields could silently change clinical display or search recall. **Next:** settle that per-field question, then either pre-truncate prose that only feeds card display, or move search matching server-side / load prose on first keystroke. **Gate:** `check:therapy-data-index` plus the therapy Playwright journeys; re-measure with `npm run verify:lighthouse`. **Stop:** do not drop a field from the catalogue payload without confirming no card renders it and no search path matches on it. Same class as #013 (route-chunk / catalogue JSON weight), different route and now measured. | PR #1915; live Web Vitals runs 31704500966 and 31704504389; codex/performance-css-delivery | 2026-07-30 | | #118 | P2 | task | Adopt the remaining visual baselines; Lighthouse now gates regressions | Lighthouse half resolved in PR #1915: authorized CI refresh run 31697669596 on current main produced all 10 route/strategy cells with one pinned HeadlessChrome/151 identity. The reviewed artifact was committed, lighthouse-budget.json enforce is true, the job no longer uses continue-on-error, merge_group coverage is restored, and pr-required now fails on a selected Lighthouse failure. The 2026-08-08 and 2026-08-13 complete baselines stayed within tolerance; the latter puts mobile LCP at 2357-2388 ms and Therapy is no longer an outlier. This relative local-production gate does not close #117 deployed-origin LCP work. Remaining #118 scope: adopt the CI-generated Linux visual snapshots and promote visual-baseline only after design-owner review and stable reruns. Stop: never use developer-machine snapshots or let a workflow update its own gate. | PR #1915; CI run 31697669596 artifact lighthouse-baseline-refresh-31697669596 | 2026-07-30 | | #150 | P2 | issue | CodeRabbit reviewed none of a full day's PRs; spending cap reached | IN FLIGHT annotation retired 2026-08-14: PR #1836 has merged, so the do-not-start note is stale and was blocking rather than protecting. The row itself is NOT code-verifiable from a container — CodeRabbit's spending cap is an account/billing state, so confirming whether the cap still suppresses reviews needs the operator's CodeRabbit dashboard. Next: check the subscription's review quota and either raise it or record the accepted coverage gap. Keeping open pending that operator read. | PRs #1404/#1430/#1444/#1445/#1479; `.coderabbit.yaml` | 2026-07-30 | -| #162 | P2 | task | Redesign Tools search results state (Compact Results Instrument) | IN FLIGHT confirmed still accurate 2026-08-14: PR #1839 is the one PR in this cluster that has NOT merged (no merge commit on origin/main; refs/pull/1839/merge still exists, which GitHub keeps only for open PRs). Every sibling in the same sweep — #1835 #1836 #1837 #1840 #1841 #1842 — has landed and their rows are archived or re-scoped. Do not start this row; it is genuinely in flight. | session 2026-07-31 mode-page design audit | 2026-07-31 | | #165 | P2 | task | Adopt a consolidated answer-home notice block — the studies exist, nothing adopts them | **Outcome:** the answer hero states its safety obligation, its scope, and its verification requirement as one block in one voice. **Detail:** `/mockups/warning-consolidation` (PR #1437) diagnoses today's three stacked notices — the APP-5 privacy warning at 11px muted, a bare `/privacy` link, and an accent-blue `ShieldCheck` capability claim at 14px semibold — and shows the hierarchy is inverted: the least important line is the loudest, and two shields with opposite meanings sit ~40px apart. Three consolidations are drawn at 1440px and 390px. Recommended: **02 Safety card** on the hero (obligation on a warning-tinted top row, everything descriptive in one grey voice below) and **01 Assurance bar** on the docked composer — the same content model at two densities, so one component with a `density` prop covers both. **This is a governance change, not just a design one:** `PrivacyInputNotice` is the single site-wide APP-5 line and renders on the answer, documents and calculators composers, so all three move together; `tests/privacy-ui.test.ts`, `tests/ui-accessibility.spec.ts` and the phone-chrome reserve coverage all assert against the current markup and must change in the same commit; and the PR will need a full `## Clinical Governance Preflight` (the mockup PR correctly did not). **Third study (before/after):** `/mockups/answer-home-proposal` draws the concrete D-direction proposal as a full hero before/after rather than an isolated notice. **Second study (words only):** `/mockups/warning-line` answers a narrower brief — no icon, border, tint or background, one line where width allows. Six variants A-F; line counts measured from the rendered DOM, not asserted. Only B (middot clauses), D (obligation + verify) and F (compressed obligation) hold one line at desktop width, and **none fit one line on a 390px phone while the pinned APP-5 sentence stays verbatim** — 46 characters of obligation plus the 27-character link exceeds the ~60 available at 11px. Recommended there: **D**, the only compliant variant that is both one line and keeps weight-only hierarchy, reached by dropping the scope claim (a capability statement already visible on the answer itself). F fits best but rewrites the pinned obligation to \|No patient-identifiable information.\| and so needs the same privacy sign-off as `#166` plus a matching `tests/privacy-ui.test.ts` update. **Status:** PR #1437 was closed unmerged on 2026-07-30 as a deliberate pause during an owner-authorized ordered merge sweep, to be reopened at its queued place; branch `claude/warning-consolidation-mockups-09jyj7` is preserved and merged onto current `main`; these follow-up rows have been renumbered on each sync because `main` kept claiming the next ids while the PR was paused; the superseded numbers are deliberately not listed, since they now belong to unrelated rows. **Next:** decide block (02 + 01) versus line (D) direction, get wording sign-off for `#166`, then implement behind one component and run `verify:phone-chrome` before `verify:ui`. | session 2026-07-30; PR #1437; `/mockups/warning-consolidation`; `/mockups/warning-line` | 2026-07-30 | | #168 | P2 | rec | Sequential issue ids force every concurrent append to conflict | DESIGNED 2026-08-14 in PR #1944 — docs/ledger-id-scheme-proposal.md. Design only, nothing implemented, so this row stays open. Recommends a ULID as the durable id with a short derived display form, the property that matters being that the display form is derived rather than stored: a clash there is a rendering fix (take one more character) rather than a renumber. UUIDv7 noted as an equally good fit. Records why timestamp-plus-slug and content hashes were rejected — the slug wants to change when a row is re-scoped, which is renumbering under another name, and a content hash is neither sortable nor stable. Migration is additive because the 314 existing sequential ids keep their numbers permanently: they are cited across the ledger, docs/branch-review-records/, AGENTS.md, the skills and the commit history, so renumbering would invalidate every citation while producing exactly the churn this row exists to end. Four steps, widening validators before allocation changes, with every current #NNN assumption enumerated by file and symbol (ledger-inbox.mjs validateRequest twice; check-outstanding-issues.mjs ID_CELL, the MARKER parse, the nextId-above-highest assertion and its padStart formatting; outstanding-issues.mjs allocator; issues-report.mjs and the issues-surface hook). Stop unchanged and now load-bearing on step ordering: do not reinstate merge=union while ids are sequential — it only becomes safe after the marker is gone. | session 2026-07-31; .gitattributes; #154/#155; PR #1524 sync | 2026-07-31 | | #169 | P2 | issue | Machine-local branches, snapshots, worktrees, and dev servers remain at risk | **CONSOLIDATED 2026-08-13 from #152, #236, and #260 before those source rows are archived by PR #1920. Outcome:** every branch, snapshot, worktree, or process that exists on only one machine remains recoverable and receives an explicit owner disposition before machine or worktree cleanup. **Original unpushed branches:** `claude/clinical-kb-design-system-333a69` was verified to contain 57 files / +4069 at tip `feat(design-system): v2 token layer, 26 components, browser-crash fix`, including `.design-sync/previews/*.tsx` absent from main. Also inspect `design-sync-db0a54`, `fable-implementation-fc937c`, `frosty-mayer-2c6167`, and `issues-133-evidence`. **Preserved WIP snapshots from #152, all unpushed, unreviewed, and unverified:** `codex/reconcile-immediate-20260730` at `748ef018f` (21 files, +395/-200 across 19 tracked, including `.github/workflows/ci.yml`, `package.json`, and `docs/scripts-index.md`); `codex/document-results-mockup-20260730` at `5dbd9f965` (8 tracked files, +13/-3, plus an untracked `document-search-results/page.tsx` mockup); `codex/chat-ledger-triage-d344` at `b7eae51a4` (`docs/outstanding-issues.md` +59/-61); and `claude/section-spy-browser-coverage` at `d949859c3` (`tests/ui-smoke.spec.ts` +51). **Wave-5 inventory from #236:** content-compare `claude/ds-v2-builder-a` and `claude/ds-v2-builder-b` with current `origin/main` because squash merges make ancestry checks unreliable; retain the associated process evidence for ports 3258 (`Database-wt-ds-v2-capture`), 3135 (`Database-wt-ds-v2-correctness`), and 3672 (`Database-wt-ds-v2-empty-state-heading`) until the owner confirms each process is no longer needed. **Stranded Sentry work from #260:** on the originating Windows machine, inspect branch `claude/cloud-pr-loop-prevention-bc052b` commits `c3c9d6a31` and `abbcdc8e9` (~389 lines across `src/sentry.*.config.ts`, `src/lib/env.ts`, `src/lib/supabase/client.tsx`, and `src/components/ui-primitives.tsx`) plus the same four uncommitted files in `.claude/worktrees/pensive-borg-6be2f0`; content-compare them with remote branches `claude/sentry-nextjs-sdk-setup-2v24q5` and `cursor/sentry-nextjs-sdk-7cee`, then record whether the work is unique, remotely preserved, or proven superseded. **Verification rule:** do not use `git rev-list` counts, three-dot diff, or ancestry alone to declare squash-merged work represented; verify the branch-added files or content against current main. **Cloud-session stop:** fresh cloud containers cannot observe the originating machine's local branches, worktrees, or processes, so never close this row from a cloud inventory that reports them absent. **Next:** complete and record each disposition from the originating machine. **Stop:** retain every listed branch, snapshot, worktree, and process record until content proof and owner disposition exist. | sessions 2026-07-30/31 and 2026-08-04/07; #152/#169/#236/#260; PR #1920 review | 2026-07-31 | @@ -161,23 +155,19 @@ removed after current-main verification; it is not missing recommended work. | #183 | P3 | task | Create Sentry metric alert for production DB span p95 > 500ms | **DEPRIORITISED 2026-08-12 (yield review against current main).** A production DB p95 latency alert for a system with one user; the alert has nobody to wake. Revisit alongside #027 when real usage exists. Still blocked 2026-08-01 closeout: SUPABASE_ACCESS_TOKEN and SENTRY_AUTH_TOKEN missing from session env; Sentry MCP OAuth can list/get alerts but has no create tool; browser hits login wall; no metric rules exist yet on clinibase-xz. Create Metric Alert: p95(span.duration), filter span.op:db, environment production, threshold >500ms, notify Active Members. Provide SENTRY_AUTH_TOKEN in session to finish via sentry alert metrics create. | session 2026-07-31 db-query-perf follow-up | 2026-07-31 | | #190 | P3 | task | X3: Finish rag.ts monolith decomposition | **DEPRIORITISED 2026-08-12 (yield review against current main).** Structural churn on the most safety-critical and most protected file in the repo, with no user-facing benefit and real behaviour-drift risk on a live-validated clinical answer path. Do the extractions opportunistically when a feature change already requires being inside a region, not as a standalone project. VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: src/lib/rag/rag.ts measures 4,362 lines — still the monolith this row describes; the decomposition has not started. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. **Outcome:** `src/lib/rag/rag.ts` is decomposed into focused modules per `docs/maturity-backlog-workorders.md` X3, with existing offline RAG contracts green. **Status:** IN PROGRESS (DocumentViewer/Dashboard extractions done; rag.ts remains). **Next:** continue safe extractions only with the RAG flag before editing protected surfaces; one verified draft PR per unit. **Stop:** no behaviour change without canary when retrieval/answer paths move. | docs/maturity-backlog-workorders.md X3; #086 | 2026-07-31 | | #191 | P3 | task | X5: ACL-migration consolidation (provider-gated) | **Outcome:** ACL-related migrations are consolidated per maturity work-order X5 without weakening owner-scope/RLS. **Next:** DB-owner approved window only; live-DB provider confirmation required before apply. **Stop:** no hosted apply from an agent session without explicit approval. | docs/maturity-backlog-workorders.md X5; #086 | 2026-07-31 | -| #192 | P3 | task | X6: Raise clinical/retrieval/answer coverage floors | **Outcome:** coverage floors for clinical, retrieval, and answer domains meet the maturity X6 targets with CI enforcing them. **Next:** set floors from current honest baselines; expand tests only where gaps are real. **Stop:** do not lower floors to pass. | docs/maturity-backlog-workorders.md X6; #086 | 2026-07-31 | | #193 | P3 | task | X7: Complete the remaining src/lib domain-directory reorg | **DEPRIORITISED 2026-08-12 (yield review against current main).** Mechanical directory moves with import-graph risk and no user-facing benefit. Same reasoning as #190: fold into work already touching the files. VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: Six domain directories exist under src/lib (extractors, observability, rag, supabase, validation, webhooks); the reorg is genuinely partial, as the row says. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. **Outcome:** remaining `src/lib` clusters sit in their domain directories per X7 follow-on to X2. **Next:** move non-protected clusters first; answer/retrieval clusters need the RAG flag. **Stop:** no drive-by behaviour edits inside moves. | docs/maturity-backlog-workorders.md X7; #086 | 2026-07-31 | | #195 | P3 | task | M1: Repo-host hardening (branch protection and required checks) | **Outcome:** GitHub branch-protection rulesets and required checks match audit §8 / maturity M1. **Next:** maintainer GitHub UI work; not a repo-file change. Record evidence in the ledger when done. **Stop:** agents must not weaken required checks. | docs/maturity-backlog-workorders.md M1; #086 | 2026-07-31 | | #206 | P2 | task | AnswerState partial_retrieval has no app-facing producer | VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: `partial_retrieval` is declared in src/lib/answer-state-types.ts:63 and handled in answer-clipboard.ts:75, but nothing in src/app or the retrieval path produces it — still no app-facing producer, as the row says. Do not synthesise it from candidate counts. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. PR-E step 0 found nothing in the client payload names which expected sources were unavailable (retrievalDiagnostics = candidate counts; conflictsOrGaps = prose). RetrievalStateBanner supports the state but PR-J adoption can only emit ready/stale_evidence/source_only. Next action: decide whether a separate RAG contract PR should add a named missing-source signal (governance preflight + RAG impact line + offline eval); until then do not synthesise the state from counts. Pinned by tests/answer-state-contract.test.ts and SPEC 13 / COMPONENTS 2. | PR-E step 0, session 2026-08-02 | 2026-08-02 | | #211 | P3 | task | Plan and start the noUncheckedIndexedAccess migration | **DEPRIORITISED 2026-08-12 (yield review against current main), and that judgment still holds** — each site is a local judgment, no open ledger row traces a defect to unchecked indexed access, and the diff conflicts with every open PR. Do it in scoped batches after the clinical and CI-trust work. This update carries that conclusion forward rather than replacing it; what has changed is that the batches now exist on paper and the count was wrong. **RE-MEASURED AND PLANNED 2026-08-14 in PR #1944.** The staged plan is docs/no-unchecked-indexed-access-migration-plan.md; the migration has NOT started and tsconfig.json is unchanged, so this row stays open and stays deprioritised. Measured against main at d47aa6d rather than reusing the 2026-08-02 figure: **1,445 errors across 269 files, up from 1,266**. The drift is itself a finding — the flag is off, so nothing stops new unchecked indexing landing, and any plan built on the stale count under-scopes. The measurement also reshapes the job in a way that supports doing it in batches: tests/ (713) plus design-scratch mockups (237) are two-thirds of the population and carry no production consequence, so the genuinely risky remainder is about 500 errors, not 1,445. Shape is 71 percent TS2532/TS18048, which a guard fixes; the 368 TS2345/TS2322 need a real decision about what the absent case means. Hot spots unchanged and confirmed: answer-verification.ts (41), rag-extractive-answer.ts (23), worker/main.ts (23), evidence.ts (19). Six stages, cheapest first, each flagged mechanical or manual with its own gate. Key constraint the plan records: noUncheckedIndexedAccess is a whole-project option and narrowing include does not isolate a directory, because TypeScript still reports errors in every transitively imported file — so the flag flips exactly once in the final PR and intermediate stages are verified by a baseline ratchet in the shape of scripts/design-system-contract-baseline.json. Stage 6 touches src/lib/rag/**, so the plan writes out the flag-before-editing, RAG impact line, and live-canary obligations. Stop unchanged: do not flip the flag on main ahead of the final stage. | session 2026-08-02 /ledger sweep — docs/review-findings-2026-08-02.md | 2026-08-02 | -| #212 | P3 | task | Replace as unknown as casts and unvalidated JSON.parse with Zod or runtime guards | RAG-surface tranche implemented via PR #1946: new src/lib/rag/rag-row-contracts.ts (assertRetrievalRows, Zod-backed, RetrievalRowShapeError) replaces 4 'as SearchResult[]' casts in rag.ts (hybrid telemetry/merge, vector-fallback, document-summary context) with runtime shape assertion -- strict on id/document_id/content (not-null in schema) and the 4 score fields (nullish), loose via z.looseObject on everything else so RPC-version column differences don't break. Assertion-only (no transform), so object identity/key order is unchanged on success; errors carry only Zod issue paths, never row content. tests/rag-retrieval-row-contract.test.ts (8 cases) plus tests/rag-imputation-contract.test.ts kept green as evidence no ranking/comparator logic moved. Deliberately untouched: query_embedding casts (a deliberate repo-wide convention satisfying Supabase-generated RPC types) and outbound Json serialization casts. Remaining #212 population (~11 as-unknown-as casts) confirmed via grep to live OUTSIDE rag.ts's import graph: src/app/api/documents/route.ts (4), src/app/api/ingestion/{batches,jobs,quality}/route.ts (6), src/app/api/jobs/route.ts (1) -- legitimate non-RAG-flagged tranche-2 candidates, not yet started. | PR #1946, session 2026-08-14 | 2026-08-02 | | #222 | P3 | task | Headers surface only partially converged in PR-J: mode-home-template and search-results-header-band untouched | VERIFIED CORRECT 2026-08-12 — re-checked against merged main and left open: Still unconverged: src/components/mode-home-template.tsx defines ModeHomeStatusNotice locally (:232) and imports neither PageHeader nor the DS EmptyState; search-results-header-band.tsx is likewise untouched. Note the adjacency — in-flight PR #1842 delegates ModeHomeStatusNotice to the DS EmptyState under #221, which is a different conversion from the PageHeader question this row asks. Re-check after #1842 merges. Builder A converged DsmPageHeader, InformationPageHeader and InformationPageBreadcrumbs onto PageHeader plus Breadcrumb, and declined two files with reasons. mode-home-template.tsx ModeHomeHero is a centred display hero on the fluid text-hero token and is the slot the in-flow phone composer sits in, so converging it onto a left-aligned PageHeader is a redesign of 13 mode homes that collides with the one-composer-per-page contract. search-results-header-band.tsx is a results spine carrying status, counts and filters, not a page-title stack, so its pin tests/search-results-header-band.dom.test.tsx remains unflipped. Both are defensible; both leave the headers surface partially adopted. Next action: decide whether either is in scope at all, or record them as permanently out of the PageHeader vocabulary. Found during PR-J adoption, 2026-08-03. | session 2026-08-03 (PR-J Wave 5, Builder A) | 2026-08-02 | | #231 | P1 | issue | Generation fallbacks no longer stick in answer cache; lithium generation quality still falls back safely | PARTIAL 2026-08-12: This PR fixes the clinically consequential stale-fallback path: every answer whose routing or degraded reason contains generation_fallback is excluded from rag_response_cache. Offline evidence: 96 focused answer-route tests and 574 RAG fixture/contract tests passed. Approved live baseline/final canaries preserved 36/36 document and content recall at 1.0 with zero per-case reciprocal-rank regressions; the final 44-case answer gate had zero citation or numeric-grounding failures. A budget extension was tested and rejected: four cache-bypassed 'Lithium dosing?' probes remained grounded, cited safe extractive fallbacks at 35-40 second candidate budgets; the decisive 40-second probe completed generation in 25.272 seconds and 27.237 seconds total with route_deadline_exceeded=false, but failed generation quality. Therefore OPENAI_ANSWER_TIMEOUT_MS and the route budget are not the current residual binding cause. INSTRUMENT NOW EXISTS 2026-08-14: the "Next: instrument" half of this row is done. Commit a3bc4da adds scripts/probe-generation-quality.ts — one cache-bypassed live answer reporting the structured generation_quality_gate_reasons, provider-backed, refusing demo mode, never caching or logging the probe. The same commit adjudicates PR #1861: superseded for phase 1, close recommended, with the numeric-retry half deferred to phase 2 pending probe evidence. So do not review #1861 as though it were the live fix, and do not re-implement the probe. Next: run scripts/probe-generation-quality.ts in an environment that has OPENAI and Supabase credentials — it is blocked in offline containers, which is why it has not been run yet — then make a separate bounded output-quality fix with an offline fixture and live canary. Stop: do not increase route/provider timeouts or cache any generation fallback. INCIDENT ADDENDUM 2026-08-14 (later the same day): rung-2 evidence was then measured live - supabase_rpc_latency_ms 31610 on a semantic query (route budget 25000 starved generation), caused by the #316 dropped trigram indexes; after their owner-approved restore, 1535 (text fast path) / 8519 (hybrid). Pre-generation latency was the binding residual cause of semantic-query source-only fallbacks in that window; evidence in docs/audit/live-drift-forensics-2026-08.md. S1 (A1 phase 2) must re-verify generation_quality_gate:* dominance on healthy latency (run the probe with node --env-file=.env.local, which the probe does not load itself) before choosing a code mitigation rung. The route-budget stop condition stands unchanged. | sessions 2026-08-14: instrument adjudication + live incident probes (owner-authorized Supabase connector) | 2026-08-04 | | #235 | P3 | task | ADOPTION.md section 7 proof shots exist for only four of the adopted surfaces | CLOSURE ATTEMPTED AND REJECTED 2026-08-14 — read this before closing again. PR #1940 queued a `done` for this row citing ADOPTION.md section 7.1's per-surface executable-evidence table; the closure was cancelled on review with the reason "executable evidence does not replace the requested desktop and phone proof shots". The cancellation is correct, and the trap is worth naming: section 7.1 opens with "This PR records executable evidence RATHER THAN committing image baselines", so the very section that looks like the evidence says in its first line that it is not. A test that proves a component is mounted is not a picture of the surface, and this row asks for the picture. IN FLIGHT note retired: PR #1842 merged, so the do-not-start warning no longer applies. The requirement is unchanged. The adoption contract asks for a proof shot per adopted surface. Wave 5 captured four - DSM header, settings rows, patient panel, answer surface - and none for the forms fold, the catalogue and docs surfaces, the headers convergence, or the empty states adopted since. Section 7 therefore reads as complete while most of the adoption is unevidenced, which matters because the proof shot is what a later reader uses to tell an intended restyle from a regression (the #229 DSM eyebrow was almost rediscovered as a defect for exactly this reason). Next action: capture the missing shots against a warmed local server (npm run ensure) and attach them to section 7. Cheap and mechanical - no gate, no provider access. Stop: this is not the visual-baseline harness (#118) - do not commit Playwright snapshot PNGs or flip that job to blocking. Stop: do not close this row on unit, DOM or contract evidence of any kind. | session 2026-08-04 (DS V2 Wave 5 close-out capture) | 2026-08-04 | -| #237 | P2 | rec | Eyeball low-confidence AccessibleTable densities at 320px before freezing Linux visual baselines | CLOSURE ATTEMPTED AND REJECTED 2026-08-14 — read this before closing again. PR #1940 queued a `done` for this row citing tests/accessible-table.dom.test.tsx:110 ("keeps the full missing-value phrase readable in the dense 320px preview"); the closure was cancelled on review with the reason "the jsdom assertion does not verify the requested real 320px browser visual pass". The cancellation is correct. jsdom does not lay out text, so a 320px assertion there proves the string is present, not that it fits, wraps or stays legible at that width — which is the whole question for a low-confidence extraction in a sparse grid. IN FLIGHT note retired: PR #1841 merged, so the do-not-start warning no longer applies. The requirement is unchanged. PR #1616 clinical MissingValue phrases increase text volume in sparse OCR grids; the contract forbids abbreviating to a dash. Next: open one real lowConfidence extraction at 320px phone width in an actual browser (npm run ensure, then a phone viewport) and accept or adjust dense preview column widths before committing Linux screenshots (#118). Stop: do not close this row on a jsdom, DOM-testing-library or snapshot-string assertion — only a real browser at 320px settles it. | PR #1616 review findings; session 2026-08-05 | 2026-08-05 | -| #238 | P2 | rec | Visual pass for Sheet portal default on settings, sidebar, and answer overlays | CLOSURE ATTEMPTED AND REJECTED 2026-08-14 — read this before closing again. PR #1940 queued a `done` for this row citing tests/sheet.dom.test.tsx and the `portal = true` default; the closure was cancelled on review with the reason "generic Sheet coverage does not verify the requested product-overlay browser journeys". The cancellation is correct. The residual risk this row exists for is ancestor-scoped CSS, `contain` and `transform` on the specific product overlays — a property of where each overlay sits in the real page, which a component-level unit test cannot see no matter how thorough it is. IN FLIGHT note retired: PR #1842 merged, so the do-not-start warning no longer applies. The requirement is unchanged. PR #1616 flips the Sheet portal default to true, moving ~10 product overlays into OverlayRoot; token inheritance is safe. Next: one visual pass in a real browser over settings-dialog, ClinicalSidebar, answer-result sheets, launcher sheet and section-nav. Stop: do not close this row on Sheet component coverage — the question is about the five host surfaces, not about Sheet. | PR #1616 review findings; session 2026-08-05 | 2026-08-05 | | #239 | P3 | rec | Manual phone rotation check for ResizeObserver-only phone chrome reserve | PR #1616 phone overlay reserve publishes only from ResizeObserver quiet-window deliveries. Desktop↔phone and late-mount recovery are covered; orientation that does not change stack height is a narrower trigger. Next: rotate a physical phone on a chrome-overlay route and confirm --phone-overlay-chrome-h updates. | PR #1616 review findings; session 2026-08-05 | 2026-08-05 | | #240 | P3 | rec | Confirm tooltip visual hard-clip asymmetry with design owner | Tooltip keeps overflow-hidden visual clamp while sr-only/aria-label retain full text. Design contract says supplementary-only. Next: design-owner confirmation that sighted users losing the clipped tail is acceptable, or allow overflow-y-auto for long clinical strings. | PR #1616 review findings; session 2026-08-05 | 2026-08-05 | | #242 | P2 | task | Commit approved Linux visual baselines and promote adoption not-committed → committed | VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: Six linux/ PNGs are committed, but the adoption manifest still carries 68 `not-committed` entries — the surfaces flip is the remaining work, as stated. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. Baselines and provenance are DONE as of PR #1729 (branch claude/ds-adopt-visual-baselines): all six linux/ PNGs committed from ubuntu artifact visual-baseline-31251091603 (main @ bc33d414e), AWAITING_BASELINE emptied, and tests/__screenshots__/linux/provenance.json written with per-candidate SHA-256 + dimensions and an approved human review. Proven by that PR's own run: visual-junit tests=9 failures=0 skipped=0, and no visual-candidates/ directory, i.e. all six compared rather than skipped. REMAINING: only the surfaces flip to baseline.status committed. Blocked on ordering, measured 2026-08-08: validateLinuxVisualBaselineSet short-circuits on declaredPaths.length===0, so declaring files activates its rule that no non-allowlisted path may change since candidateSourceHead — and PR #1729 necessarily changed tests/design-system-adoption.test.ts, whose initialiseCandidateRepository seeded fixtures from the LIVE spec and so failed the moment AWAITING_BASELINE emptied. The two cannot land together. Next: after #1729 merges, re-capture candidates from a main run that already contains that fixture fix, then flip the surfaces against that head. Note this does not affect whether pixels compare — Playwright compares because the goldens exist on disk. | PR #1616 review findings; session 2026-08-05 | 2026-08-05 | | #248 | P2 | issue | Investigate why 20260705180000 search-health indexes were missing on live despite applied history | APPEND 2026-08-13: the prior closure is withdrawn. Repository and live-drift evidence establishes that 20260705180000_reconcile_search_health_indexes.sql is recorded as applied while documents_title_trgm_idx and document_chunks_content_trgm_idx are missing on live. Supabase transaction semantics exclude a persisted partial migration, but the present record does not distinguish skipped DDL/history repair from indexes created and later dropped. In an approved read-only window, query supabase_migrations.schema_migrations for the 20260705180000 statements fingerprint and inspect the relevant audit/history evidence; retain both hypotheses until that evidence establishes the cause. Separately, scheduled check:drift did detect the missing indexes, but red runs were not routed. | PR #1614 review / session 2026-08-05 (renumbered on main merge) | 2026-08-05 | | #258 | P2 | rec | The PR-handoff stop rule is enforced for Claude Code only; Codex and Cursor get prose with no gate | GAP RECORDED 2026-08-14 in PR #1944 — docs/pr-handoff-stop-cross-agent-gap.md. This is the row's own stated fallback ("If no mechanism exists at all, record that explicitly here so the gap is a known limit rather than an open task"), so the row stays open but is no longer unexamined. Checked, not assumed: .claude/settings.json is read only by Claude Code; plugins/clinical-kb/.codex-plugin/plugin.json declares name/version/description/author/repository/keywords/skills and an interface block with NO hook, event, or pre-tool-interception field, shipping exactly one skill; .cursor/ holds settings.json (plugin enablement only), mcp.json, agents/ and skills/ with no deny path. So the cheapest-first option the row proposed is currently unavailable in both tools. Worth noting because it sharpens the cost: .cursor/agents/pr-babysit.md exists, meaning Cursor ships a documented agent for exactly the PR-following behaviour this rule restricts, with nothing bounding it. The doc records the Claude Code mechanism in enough detail to reimplement (session-scoped marker under the absolute git dir, fail-open on an unidentifiable session id, never pruning a sibling's marker, post-mode scanning only the request half so a command that merely prints a PR URL cannot arm it, and the CLAUDE_ALLOW_PR_FOLLOW=1 prefix unlock that a mention alone cannot trigger), plus the three questions any parity mechanism must answer. It is explicit that the wrapper fallback is advisory only — it cannot touch the MCP-connector or loop-machinery classes, so it makes a violation detectable after the fact rather than prevented. Next: re-check the Codex and Cursor manifests when either ships hook support; close only when a mechanism exists or the limit is accepted deliberately. Stop unchanged: do not weaken the Claude Code hook for symmetry, and do not keep a second copy of the deny list. | PR #1649; .claude/hooks/pr-handoff-stop.sh; .claude/settings.json; AGENTS.md "Stop when the pull request is open"; session 2026-08-07 | 2026-08-07 | -| #265 | P2 | task | DS Track A6: move design-system gates 2, 4, 7 and 8 from partial to blocking | GATE 2 CLOSED 2026-08-15 by PR #1984; gates 7 and 8 remain, and gate 8 is materially larger than this row said. Corrections first, because two figures here were stale: the debt for gate 8 is 25 conflicts across 12 files, not 27 across 15 — PR #1942 paid three files down without updating the prose, and both GATES.md section 3 and scripts/design-system-contract-baseline.json already said 25, which outranks the prose per docs/design-system/README.md. GATES.md carries 9 implemented-partial rows, not 10; the tenth grep hit is the legend. And #293 is resolved: commit 885c613 (PR #1962) IS on main — PR #1977 is already open to close that row, so do not queue another. GATE 2, what closed it: the row named two gaps, the rendered enumeration and the fixed-height h-10 case. The first landed in #1962. The second is closed by a new interactiveTapFloorDeclarations metric in check:design-system-contract — an interactive element (a, button, input, select, summary, textarea) declaring its OWN unprefixed min-h-* below the 48px token, ratcheted at 43 across 17 files with per-path pins; comparable arbitrary lengths and mutually exclusive/composed branches are evaluated independently, so a new sub-floor control anywhere in src/** fails today. Mutation-verified: lowering one shortlist button to min-h-9 produced both the total and matching per-path regression. Scoped to min-h-* and NOT h-*/size-* on purpose: a short h-4 on an interactive element is routinely the visible box of a control whose hit area belongs to a tap-sized wrapper (SelectionCheckbox in differentials-home.tsx, whose label ui-smoke asserts still meets the floor), and flagging those would pad the baseline with non-defects — the exact section 5 failure mode. One remaining limit recorded rather than hidden: the walker sees intrinsic lowercase tags only, so a floor on or another component wrapper is invisible to it (the pre-existing legacyTapClasses check shares that blind spot). The 43 recorded sites are real debt still owed. Also tightened legacyShadowAliases 119 to its measured 118 — one unit of stale slack found while working #163. GATE 8 — STOPPED DELIBERATELY, and the reason matters for whoever picks it up. The 25 conflicts are not 25 mechanical one-line edits. Inspected every site: the large majority are focus:ring-4 focus indicators co-existing with a border on inputs, selects and textareas (master-search-header 5, formulation-builder 3, and singles across formulation-compare, dashboard-nav, favourites, specifiers, DocumentTagCloud, ui-primitives), and only pwa-lifecycle's shared cardClassName is the decorative persistent double edge the rule was written for — one recipe counted 5 times. Retiring the focus-ring class means restyling focus indicators from ring to outline across roughly 17 controls, which is an accessibility-visible change needing focus-state proof in a browser. Three further constraints: this container cannot run Chromium (ships chromium-1194 against a pinned 1234, #255/#312) so that proof was unavailable; 8 of the 25 sit in files that open PRs #1976, #1982 and #1983 are editing; and pinning at zero without first widening the onePixelShadowSpreads property filter (design-system-contract-utils.mjs:1320 matches --e[0-4] and --shadow-* but not --glow-primary, --glow-soft or --ring-hairline, all of which carry 0 0 0 1px) would be a partly false close. Next for gate 8: decide the focus-indicator question first as a design-system ruling (ring vs outline for focus), then widen the spread filter, then retire and pin — with browser focus proof. GATE 7 untouched, unchanged from this row's description: no child/parent elevation check exists and no shared render-tree traversal helper exists to build one on; every spec inlines its own page.evaluate walk. The template is #1962's determinism scaffolding plus a computed-box-shadow to --e0..--e4 tier lookup built with the probe technique at ui-style-contract.spec.ts:265-289. | session 2026-08-15; PR #1984; GATES.md; scripts/design-system-contract-baseline.json | 2026-08-07 | +| #265 | P2 | task | DS Track A6: move design-system gates 2, 4, 7 and 8 from partial to blocking | CORRECTION QUEUED 2026-08-15 during PR #1987 review. Gate 2 is closed for new use, but the just-reconciled detail inherited stale measurements from PR #1984's pre-sync tree. On the exact current tree after PRs #1982-#1986, check:design-system-contract measures interactiveTapFloorDeclarations=40 across 16 production files, not 43 across 17, and edgeOwnershipConflicts=19 across 10 files, not 25 across 12. The strengthened tap-floor parser removed three old false-positive counts from calculators/search-page.tsx; favourites-library-nav and pwa-lifecycle no longer contribute edge conflicts after the merged UI work. The baseline and GATES.md still permit/report the older 43 and 25 values, so they carry three and six units of stale ratchet slack respectively and need a focused tightening follow-up. The detector still correctly evaluates comparable arbitrary lengths and reachable conditional/composed branches, and component-wrapper tags such as Link remain its known blind spot. Gate 7 remains open: add a shared deterministic render-tree traversal plus child/parent elevation-tier check. Gate 8 remains open: decide ring-versus-outline focus ownership, widen onePixelShadowSpreads to all relevant token families, then retire conflicts with browser focus proof. Do not edit the applied 8c133c4e request in place; it is immutable audit history. | PR #1987 review; exact tree after #1986; check:design-system-contract; scripts/design-system-contract-baseline.json; docs/design-system/GATES.md | 2026-08-07 | | #266 | P3 | task | DS Track B1: adopt the 23 unadopted components demand-driven, never as a race to 53/53 | **DEPRIORITISED 2026-08-12 (yield review against current main).** Adoption counting toward 53/53 while a clinical P1 is open. The row's own title says never as a race to 53/53; the queue has been running the race anyway. Demand-driven means it activates when a surface needs a component, not on a schedule. COUNTS RE-MEASURED 2026-08-12 from docs/design-system/adoption-manifest.json on merged main: **54 registered, 31 adopted, 23 UNADOPTED**. (This supersedes the 2026-08-08 figures of 53/30/23, which a main-merge briefly restored over this correction.) The total held at 23 but the membership moved — DisclosureGroup joined the adopted set, and the newly built ErrorState joined the unadopted set; ErrorState's enforcement is closed (archived #298) but its adoption is still open under #299. Today's 23: AnswerFooter, Checkbox, Citation, CitationList, ConfirmDialog, Disclosure, DoseLine, DownloadLink, ErrorState, ErrorSummary, ExternalTextLink, FieldError, FieldHint, LinkAction, Pagination, Progress, RadioGroup, SearchField, StageList, Tabs, TextLink, ToastRegion, Tooltip. Approach unchanged and still correct: demand-driven adoption — pick a surface and let it pull, the way AccessibleTable pulled Button and the answer surface pulled AnswerCard (#216) — never a race to 54/54. Forms remain the largest single tranche: FieldError, FieldHint, ErrorSummary, SearchField, Checkbox and RadioGroup land together on one form conversion. Do not stub a component to move the count. Regenerate with npm run design-system:adoption:update AND npm run design-system:design-sync:update; both manifests are generated, never hand-edited. | session 2026-08-07 — design-system HANDOVER-2026-08-07 Track A1 handoff (PR #1678) | 2026-08-07 | | #267 | P3 | task | DS Track B2: AnswerFooter and DoseLine need a provenance/dose payload the answer surface does not produce | **DEPRIORITISED 2026-08-12 (yield review against current main).** Blocked on a provenance/dose payload the answer surface does not emit, which is backend work nobody has scoped. Cannot start. VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: Neither AnswerFooter nor DoseLine has a product importer; the provenance/dose payload the answer surface would need still does not exist. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. Backend-shaped work, not a component swap: the two components cannot be adopted until the answer surface emits the provenance and dose data they render. Do not stub one to make the adoption count look better. Sequence after the payload exists, then adopt via the Track B1 demand-driven route. | session 2026-08-07 — design-system HANDOVER-2026-08-07 Track A1 handoff (PR #1678) | 2026-08-07 | | #268 | P3 | task | DS Track B3: move the 19 genuine bare-dash sites onto MissingValue | **DEPRIORITISED 2026-08-12 (yield review against current main).** 19 bare-dash sites with no reported clinical misreading. Cosmetic consistency on a prototype with an open P1. VERIFIED CORRECT 2026-08-12 — re-checked against merged main during the full ledger sweep and left unchanged: MissingValue is imported in 5 component files; the bare-dash conversion is partial. The ~5 calculator 'not started' sites stay permanently, per this row's own stop rule. This stamp exists so a later reader can tell "checked and still true" from "never looked at"; the two were indistinguishable before. Therapy-compass getters, specifier sourceFamily, favourites counts when untrusted. Leave the roughly 5 calculator 'derived.started ? score : dash' sites PERMANENTLY — 'not started' is not a missing clinical value, MissingValueReason has no member for it, and converting them would render 'Not recorded' for a score the clinician simply has not entered. | session 2026-08-07 — design-system HANDOVER-2026-08-07 Track A1 handoff (PR #1678) | 2026-08-07 | @@ -194,23 +184,21 @@ removed after current-main verification; it is not missing recommended work. | #309 | P2 | task | Facet groups of 6-20 options render as chips, not the dense list docs/filter-contract.md section 5 requires | Attempted 2026-08-14: an implementation task for chips-for-6-20 was stopped before any code was written, because it directly contradicts this row's own current, still-open text, which requires a full-width DENSE LIST (right-aligned count column, group headings) for the 6-20 band, and explicitly says chips-for-6-20 does not satisfy this row. Confirmed chips-for-6-20 is ALREADY the live behaviour (dense = facetGroups.length > 3 \|\| totalFacetOptions > 20 in result-filter-control.tsx), and that closing this row on that basis was already tried once and explicitly reverted (PR #1925, 'correct #309 to partially delivered'). No code changed, no PR opened. Needs a product/design decision between: (1) build the genuine full-width dense-list renderer plus the nine-option DOM assertion this row asks for, or (2) formally amend docs/filter-contract.md section 5 to deliberately drop the middle band with reviewer sign-off -- different from what already happened (a silent merge-conflict resolution the row says didn't count). | session 2026-08-14, agent stop per contract contradiction | 2026-08-12 | | #311 | P3 | task | Promote the derived ledger loss-detector into scripts/ — it has now earned its place twice | During the 2026-08-12 sweep, two main-merges silently reverted edits to `docs/outstanding-issues.md`, including the ENTIRE #293 refutation (a `grep sm:min-h-0` returned 0; the text survived only in commit a6bfc6f). It went unnoticed because the recovery script was HAND-ENUMERATED — it listed 15 archives and 8 updates from one commit and could therefore only restore what the author remembered. The replacement is derived rather than listed: read every row id this branch has ever stamped out of `git rev-list ..HEAD` plus `git show :docs/outstanding-issues.md`, then assert each of those ids that is still OPEN carries its stamp text, and exit non-zero listing any that lost it. It has now proved itself twice — it caught the intentional #262 divergence (main's version was newer than the branch's, correctly left alone) and would have caught the #293 loss the hand-written list missed. The plan that created it said it should stay a scratch script 'unless it proves useful more than once'; that condition is met. Next: port it to scripts/ (suggested `check-ledger-stamp-retention.mjs`), generalise the stamp token from the hard-coded 2026-08-12 date to a `--since` or marker argument, add a self-test in the style of the other ledger scripts, and document it beside `ledger:dedupe` for use after any main sync that touches the ledger. Stop: do NOT wire it into verify:cheap or CI — it is a branch-local safety net for a human or agent mid-sweep, and it has no meaning on a branch that has not stamped rows. Related: #156 and #168, which track the id-allocation race that produces these merges in the first place. | session 2026-08-12 ledger sweep; scratch loss-check.mjs; #293 restoration from a6bfc6f | 2026-08-12 | | #312 | P3 | issue | check:playwright-browser-revision reporting OK does NOT mean browsers are installed — and installing the matching revision is a cheap first option | Progress 2026-08-15: PR #1965 landed on main (commit 3ec6116) and closes the false-OK gap for the pinned Chromium revision by resolving the effective cache and requiring a launchable binary. Keep this issue open: the unscoped test:e2e and release matrix also require Firefox and WebKit, and the check does not yet report whether their locked revisions are installed. Next: enumerate required and installed revisions for all browser families, with a Chromium-only-cache regression; project-scoped --project=chromium runs may continue to require Chromium alone. | session 2026-08-12; scripts/playwright-browser-preflight.mjs:127-152; scripts/run-playwright.mjs:50-53; #290 close-out; archived #255 | 2026-08-12 | -| #314 | P2 | issue | Ship compact compressed registry projections and verify live transfer | Next: land the existing view=summary/search and gzip implementation, deploy it, then verify /api/registry/records on the exact deployment SHA returns counts-only home responses and compressed compact search responses. Why: the live full payloads measured on 2026-08-13 were 482786 bytes for Forms and 1096689 bytes for Services and were downloaded by count/search-only consumers without Content-Encoding. The local projections reduce raw search data by about 91.3% and 82.0%, with gzip responses about 4.9 KB and 27.3 KB. Context: latency and Sentry review. Owner: assistant. Confidence: high. Depends on: #013 and #016. Gate: focused registry/consumer tests, production build and bundle budget, then post-deploy headers/bytes and live LCP rerun. Stop: do not close from local-only payload measurements or deploy without explicit authorization. | session 2026-08-13 latency review; src/app/api/registry/records/route.ts | 2026-08-13 | +| #314 | P2 | issue | Ship compact compressed registry projections and verify live transfer | UPDATE 2026-08-16: the provider-free implementation is now on current main via commit ca788d41e: view=summary/search projections and gzip are shipped in repository code. Remaining scope is external-only: deploy an exact authorized SHA, verify /api/registry/records headers and transfer bytes on that deployment, then rerun live LCP. Stop: do not close from local payload measurements, and do not deploy or query production without explicit target authorization. | Commit ca788d41e on current main; original latency audit evidence | 2026-08-13 | | #315 | P3 | rec | If the ui-smoke scroll-hide flake (archived #290) recurs, start from the reporter-stranding mechanism — and treat the old regression window as unconfirmed | Independent verification on 2026-08-13 (second session, fresh cloud container, pinned Chromium 1234 installed per #312) measured the archived #290 flake at BOTH ends of its recorded window and corrects the archive's causal story: the bad SHA 9ab3b73ad itself passed 16 recorded executions — reproducer isolated --repeat-each=5 (5 passed, ~1.0s each), one full tests/ui-smoke.spec.ts --project=chromium run (98 tests passed, 2.5m, 0 flaky), and reproducer x10 under deliberate CPU contention (6 busy-loop processes on 4 cores, run times 1.2-1.5s: 10 passed). Current main a76f280 also 5/5. So the recovery was NOT drift — the exact commit that measured 2/5-3/5 failures passes cleanly here — and the e8adde1b9..9ab3b73a window is unconfirmed; the failure was specific to the original machine's environment/load profile. Recorded as a comment on PR #1884 (issuecomment-5272932999). On recurrence, do not re-bisect first: test the stranding mechanism. computeScrollHideUpdate (src/components/clinical-dashboard/use-hide-on-scroll.ts) re-evaluates only on scroll/resize events, and its viewportHeightChanged / maxOffset-range-change guards deliberately zero accumulated down-travel (contract-asserted in tests/use-hide-on-scroll.test.ts) — so geometry churn consuming the final steps of a gesture strands the not-hidden state permanently until the next event, matching the recorded ~11.5s toHaveAttribute timeout signature (the assertion DOES auto-retry for 10s; the attribute genuinely never flips). Fastest confirmation: a diagnostic page.on('console') trace logging which guard fires per evaluation. The window itself was one PR (#1744 mode-routing, true merge a503c22) whose net diff touched no scroll-hide code — content-bisect axes, if ever needed: tests/ vs src/ split, use-home-mode-seed/use-last-app-mode neutralized, prefetchModeDestination reverted, positional heading click restored to a settle wait. Stop: any guard change is a behaviour change to protected phone chrome — needs a failing trace first, never speculatively; do not weaken the assertion or tap targets. | session 2026-08-13; PR #1884 comment; archived #290; #312 | 2026-08-13 | -| #316 | P1 | issue | Live DB has 20 currently missing repo-defined indexes and 10 retrieval RPC bodies diverge; weekly live-drift has been red since 2026-07-26 with no routing | Combined 2026-08-14 update, superseding the two partial requests cancelled in this same batch. PHASE 0 CLOSED including the forced-dispatch proof its definition of done required: live-drift dispatched on main (Actions run 31813064485) failed at the drift step, the always() capture step still ran, the migration-history step correctly skipped, and the separate drift-routing job then created issue #1963 "Live drift check failing" carrying the label, run URL, job result, trigger and the full findings block. Routing is now also covered offline by tests/live-drift-workflow.test.ts, mutation-verified. INCIDENT REPAIR, owner-approved in-session: the two retrieval-critical indexes documents_title_trgm_idx and document_chunks_content_trgm_idx were restored with CREATE INDEX CONCURRENTLY plus ANALYZE, both indisvalid and indisready at 648 kB and 68 MB, re-verified afterwards by an independent read-only query. Before and after supabase_rpc_latency_ms 31610 to 1535 on the text fast path and 8519 hybrid, with match_document_chunks_text_v2 at 14 ms. No repo schema change was needed because the definitions were already codified. CORRECTED FIGURES measured 2026-08-14, superseding the 2026-08-09 numbers this row was opened with: 10 match_* def_hash mismatches (unchanged), 20 missing_live indexes rather than 21, and the same 2 unexpected_live. ATTRIBUTION STILL OPEN: migration 20260705180000 recorded 14 executed statements so it was not mark-applied, and the 20260804110240 guard validates four other indexes and never checks this pair, so it gives no existence bound for 2026-08-04. The drop window is therefore 2026-07-05 to 2026-08-02 and the dashboard audit-history pairing remains owner action; #248 stays open. NEXT: Phase 3 RPC reconciliation before Phase 4, per the plan's ordering that the change which can alter clinical answers precedes the ones that only speed them up. Evidence: docs/audit/live-drift-forensics-2026-08.md. | PR #1968 review against docs/audit/live-drift-forensics-2026-08.md, 2026-08-15 | 2026-08-13 | +| #316 | P1 | issue | Live DB has 20 currently missing repo-defined indexes and 10 retrieval RPC bodies diverge; weekly live-drift has been red since 2026-07-26 with no routing | Combined 2026-08-14 update, superseding the two partial requests cancelled in this same batch. PHASE 0 CLOSED including the forced-dispatch proof its definition of done required: live-drift dispatched on main (Actions run 31813064485) failed at the drift step, the always() capture step still ran, the migration-history step correctly skipped, and the separate drift-routing job then created issue #1963 "Live drift check failing" carrying the label, run URL, job result, trigger and the full findings block. Routing is now also covered offline by tests/live-drift-workflow.test.ts, mutation-verified. INCIDENT REPAIR, owner-approved in-session: the two retrieval-critical indexes documents_title_trgm_idx and document_chunks_content_trgm_idx were restored with CREATE INDEX CONCURRENTLY plus ANALYZE, both indisvalid and indisready at 648 kB and 68 MB, re-verified afterwards by an independent read-only query. Before and after supabase_rpc_latency_ms 31610 to 1535 on the text fast path and 8519 hybrid, with match_document_chunks_text_v2 at 14 ms. No repo schema change was needed because the definitions were already codified. CORRECTED FIGURES measured 2026-08-14, superseding the 2026-08-09 numbers this row was opened with: 10 match_* def_hash mismatches (unchanged), 20 missing_live indexes rather than 21, and the same 2 unexpected_live. ATTRIBUTION STILL OPEN: migration 20260705180000 recorded 14 executed statements so it was not mark-applied, and the 20260804110240 guard validates four other indexes and never checks this pair, so it gives no existence bound for 2026-08-04. The drop window is therefore 2026-07-05 to 2026-08-02 and the dashboard audit-history pairing remains owner action; #248 stays open. NEXT: Phase 3 RPC reconciliation before Phase 4, per the plan's ordering that the change which can alter clinical answers precedes the ones that only speed them up. Evidence: docs/audit/live-drift-forensics-2026-08.md. OFFLINE ANALYSIS 2026-08-17 (no live call made; Supabase MCP was unauthenticated in that session). THIS ROW'S 'NEXT: Phase 3' IS NOT EXECUTABLE AS WRITTEN - correcting it is the point of this update. Phase 3's prerequisite in docs/database-remediation-playbook.md is the Phase 1 dossier with '10 diffs classified', and docs/audit/live-drift-forensics-2026-08.md section 1.2 records all ten as UNCLASSIFIED with 'per-function diff hunks still pending'. Phase 3's own pasted prompt says any UNCLASSIFIED entry stays untouched and is escalated to the owner. So Phase 3 currently has ZERO executable entries and a production window would accomplish nothing. THE ACTUAL NEXT STEP IS A READ-ONLY WINDOW to finish Phase 1.2, which is far cheaper and lower risk than the production window this row's wording implies. Do not request a production window first. TESTABLE HYPOTHESIS PREPARED OFFLINE that should shorten that read-only window: def_hash is computed in migration 20260706200000 as md5 of pg_get_functiondef with block comments, line comments and all whitespace stripped - it does NOT strip SET attributes, which pg_get_functiondef renders. Migration 20260724000000_optimize_rpc_work_mem.sql applies ALTER FUNCTION ... SET work_mem = '64MB' to exactly EIGHT of the ten UNCLASSIFIED functions (match_document_chunks_hybrid, match_document_embedding_fields_hybrid, match_document_index_units_hybrid, match_document_memory_cards_hybrid, match_document_memory_cards_hybrid_v2, match_document_chunks_text, match_document_lookup_chunks_text, match_document_table_facts_text). So one read-only query - does live carry SET work_mem on those eight - plausibly classifies 8 of 10 in a single step, and if live lacks it they are repo-ahead but NON-BEHAVIOURAL for answer content (a planner memory setting affects latency, not ordering or content), so they would not need the full eval-canary pair Phase 3's repo-ahead rule assumes. This is a HYPOTHESIS, not a finding: live state was never read. The two remaining outliers, match_document_chunks_text_v2 and match_document_index_units_hybrid_v2, are absent from that migration and need their own diffs; their canonical bodies are in 20260717160000_optimize_owner_public_retrieval.sql, 20260713020000_owner_plus_public_retrieval.sql and 20260717162000_bound_versioned_retrieval_match_count.sql. DISCARDED REASONING, recorded so nobody repeats it: 'the drift manifest never mentions work_mem' is NOT evidence for or against the hypothesis, because supabase/drift-manifest.json stores only signature, def_hash and acl per function and never stores body text at all. Flag before editing: this whole surface is protected RAG retrieval, so any actual change needs the RAG-surface flag, the PR RAG impact line, and per-RPC approval as the playbook already specifies. | PR #1968 review against docs/audit/live-drift-forensics-2026-08.md, 2026-08-15 | 2026-08-13 | | #317 | P2 | task | Verify registry-backed service records preserve facet metadata | #1878 introduced the services filter-contract tree and #1882 later merged the identical tree, so no merge-conflict audit is required. Current main uses ServiceRecord.catalogPayload.tags and fixture coverage verifies 219 records. Add focused offline tests that recordToRow and rowToServiceRecord preserve all six tag dimensions and degrade safely when payloads are malformed or absent. Do not add a second facets carrier unless a failing test proves the current contract inadequate. | PR #1921 review; #1878/#1882 tree comparison; service-facets.ts; registry-records.ts | 2026-08-13 | -| #318 | P1 | task | The medication interaction lexicon has never been clinically reviewed and its sign-off block is empty | docs/medication-interaction-lexicon-review.md is generated by npm run medications:lexicon-report and expands every lexicon term to the catalogue drugs it resolves to, with how many CRITICAL/HIGH rows depend on it, sorted by severe usage. It is marked UNREVIEWED and its sign-off table is unfilled, so every red and amber drug-drug interaction alert is currently an unvalidated mapping over source-backed text. The wording shown to a clinician is always verbatim catalogue prose; what is unreviewed is which drugs a phrase like 'NSAIDs' or 'CNS depressants' was taken to mean. The sheet has already produced three defects on generation alone (ARB matching Carbapenem across 16 CRITICAL/HIGH rows; two divergent Warfarin records; lithium unreachable from eight HIGH rows), which is a fair indication of what reading it would still find. Next: a clinician reads the term table top-down (it is sorted so the top ten terms carry most of the severe usage) and fills in the sign-off block. Stop: do not treat check:medication-lexicon-report passing as review - that check only proves the sheet describes the current lexicon, not that the mappings are correct. | PR #1923; docs/medication-interaction-lexicon-review.md; docs/samd-classification-medication-considerations.md | 2026-08-13 | +| #318 | P1 | task | The medication interaction lexicon has never been clinically reviewed and its sign-off block is empty | docs/medication-interaction-lexicon-review.md is generated by npm run medications:lexicon-report and expands every lexicon term to the catalogue drugs it resolves to, with how many CRITICAL/HIGH rows depend on it, sorted by severe usage. It is marked UNREVIEWED and its sign-off table is unfilled, so every red and amber drug-drug interaction alert is currently an unvalidated mapping over source-backed text. The wording shown to a clinician is always verbatim catalogue prose; what is unreviewed is which drugs a phrase like 'NSAIDs' or 'CNS depressants' was taken to mean. Next: a clinician reads the term table top-down and fills in the sign-off block. Stop: do not treat check:medication-lexicon-report passing as review - that check only proves the sheet describes the current lexicon, not that the mappings are correct. WORKLIST PREPARED 2026-08-15 (PR #1991): docs/medication-lexicon-review-worklist.md gives the top ten terms by severe usage (236 of 390 severe firings, 61 percent) with resolved drug sets and six prioritised questions. TWO OF THE THREE DEFECTS THIS ROW CITES WERE ALREADY CLOSED: the ARB/Carbapenem substring match is fixed and guarded, and lithium is reachable (9 rows / 9 severe). The divergent Warfarin pair remains and is worse than stated - warfarin-vka and warfarin-anticoagulant carry 3 interaction rows each with ZERO in common, so which record is opened changes which warnings appear. TWO FIXES LANDED 2026-08-17, owner-approved, both mechanical rather than clinical. (1) DEAD SLUG: the tcas selector listed slug 'dothiepin' but the catalogue keys the drug as 'dosulepin' (same drug, current INN), so the slug matched zero records and Dosulepin - whose own record flags Toxicity in OD FATAL and Anticholinergic HIGH - fired none of the term's 20 CRITICAL/HIGH rows. Fixed; restoring the author's evident intent, corroborated by the catalogue already filing it subclass TCA. Measured effect after regenerating data/medication-interaction-index.json: 22 rows now name dosulepin as a counterparty, 20 of them CRITICAL/HIGH, up from 0 via this term; aggregate resolution is unchanged (523 rows, 362 resolved, 161 unresolved, 423 with a catalogue target) because those rows already resolved through other TCAs, so this widens counterparties inside already-resolved rows rather than resolving new ones. Durable guard added: the coverage test now fails on ANY selector slug or denySlug that resolves to no catalogue record. The pre-existing test only required a TERM to resolve to some drug, so tcas stayed green on five of its six slugs - that is exactly how this shipped. (2) THE REVIEW INSTRUMENT'S TWO BLIND SPOTS: missedClassMembers() in scripts/build-medication-lexicon-report.ts skipped any surface stem shorter than four characters, which made the check unable to fire at all for tcas and arbs (ppis was rescued by its long surface 'proton pump inhibitors'), and it read only class and subclass, never tag. So the sheet's printed 'Checks that ran and found nothing' line was false for two terms - a printed clean result that could not have found anything is worse than no line, because it retires the question. Both closed: the floor is now 3, the shortest stem any real surface produces, and the haystack includes tag. The sheet now raises the Celecoxib/Parecoxib coxib gap itself (2 flagged, up from 1). Design note recorded because the first attempt was wrong: the fix originally matched short acronyms as whole tokens, and mutation testing showed that branch did no protective work - the leading word boundary already stops 'arb' reaching inside 'Carbapenem' - while it would newly MISS a subclass spelled 'TCAs', a regression in the dangerous direction. It is a plain prefix match, pinned by a pluralised-subclass test. STILL OPEN AND STILL YOURS: the sign-off block is untouched and the sheet is still UNREVIEWED, which is the only thing that closes this row. Five clinical questions remain with their mappings deliberately unchanged - nsaids excluding Celecoxib/Parecoxib across 38 severe rows (now auto-flagged); maois excluding Moclobemide across 17 severe rows, which the sheet still CANNOT surface because Moclobemide's tag is also RIMA and RIMA/MAOI are synonyms in pharmacology but unrelated as strings; opioids including Loperamide across 35 severe rows in the false-alert direction; acei and arbs resolving to one drug each, which is catalogue coverage rather than a narrow selector (ramipril, lisinopril, irbesartan, telmisartan, valsartan are absent from the catalogue entirely); and anticoagulants including three antiplatelets while deliberately excluding Aspirin on identical class metadata. Also worth its own row: src/lib/medication-interaction-lexicon.ts alone classifies clinicalRisk FALSE under classifyPullRequestFiles, and only the generated data/medication-interaction-index.json makes a lexicon PR clinical-risk - so a lexicon edit that changes which drugs a CRITICAL phrase resolves to would skip the governance preflight if the index were not regenerated in the same PR. | PR #1923; docs/medication-interaction-lexicon-review.md; docs/samd-classification-medication-considerations.md | 2026-08-13 | | #320 | P3 | task | Crop-to-page overlay remains unbuilt; bbox already reaches viewer state at runtime but is untyped, unvalidated, and unused | **Outcome:** selecting an indexed table or diagram can highlight its region on the PDF page, or the capability is deliberately retired — either way it stops living only in a plan document. **Detail:** this is the one Phase 3 capability never built (docs/plans/document-viewer-redesign-plan.md, Phase 3 table, 'Out of scope'). It had no ledger row until now, which is how work disappears between sessions: the plan doc marks it out of scope and nothing in durable memory says it remains owed. **The data path is partially live, not dropped.** src/lib/document-detail.ts SELECTs bbox alongside the other image columns, and withImageTableMetadata spreads every selected field except metadata. bbox therefore survives the runtime response and reaches DocumentViewer's image state. The gap is static and behavioural: DocumentDetailImage in src/lib/document-detail-contract.ts does not declare bbox, ImageRow in src/components/document-viewer/types.ts aliases that contract, no normalisation validates the stored value, and no viewer code renders it. Verified against exact PR head 2ac0f48a820be62947112efbb5d0845a702dad8e on 2026-08-13. **Shape of the work, in order:** (1) establish the ingestion coordinate space and stored shape, add a normalised bbox field to DocumentDetailImage, and add a focused loader or route-serialization test proving bbox survives with the promised shape. Do not change the selected-field mapping unless that test demonstrates an actual loss. (2) Only then draw the highlight over the rendered page when a figure is selected, accounting for the virtualized page column, the per-page raster scale from resolveViewportScale, and rotation. **Why it was scoped out rather than overlooked:** the contract and normalisation work has a wider blast radius than the component-only Phase 3 diff, and crop geometry quality from ingestion is separate debt — the redesign plan's residual-risk section says not to block viewer UX on perfect crops. **Stop:** do not land the typed-contract and normalisation half inside a viewer-only PR; it changes what the document-detail API promises and needs its own review and governance preflight. Do not render raw, unvalidated bbox values — a highlight over the wrong region of a clinical source is worse than no highlight. | session 2026-08-13 document-viewer remaining-work inventory; docs/plans/document-viewer-redesign-plan.md Phase 3 table; src/lib/document-detail.ts bbox projection | 2026-08-13 | | #321 | P3 | task | Four follow-up groups cover nine controls after #291 | Six controls in the differential comparison page stay coupled to its planned rewrite and pinned density test. The filmstrip Page unknown control is a later mechanical change. DocumentViewer needs its persistent access reason split from transient loading before classification. The pin-limit control remains a capacity-state judgement. These are four source groups and nine controls, not four controls. | PR #1778 body; verified against main 2d27039 | 2026-08-14 | -| #322 | P2 | issue | Two catalogue records are both named Warfarin and share no interaction rows, so which one a clinician opens changes the warnings | data/medications-snapshot.json holds warfarin-vka and warfarin-anticoagulant, both displayed as 'Warfarin', both class Anticoagulant / subclass Vitamin K Antagonist. They carry three interaction rows each with ZERO in common, so the alerts a clinician sees depend on which record they happened to open, and nothing on screen distinguishes them. A lexicon class term resolves to both. This is a catalogue DATA defect, not a lexicon fault - merging, deleting one, or relabelling them is a clinical content decision, which is why it is reported rather than patched. Surfaced automatically by duplicateCatalogueNames in scripts/build-medication-lexicon-report.ts, which compares the row sets and states the divergence rather than asking about it, and pinned by a test in tests/medication-interaction-lexicon-coverage.test.ts that goes red when the records are reconciled so the flag can be retired with it. Next: a named clinical owner decides the disposition. Stop: do not de-duplicate by display name in the report or the UI - that hides the divergence rather than resolving it. | PR #1923; docs/medication-interaction-lexicon-review.md flag section; tests/medication-interaction-lexicon-coverage.test.ts | 2026-08-13 | +| #322 | P1 | issue | Two catalogue records are both named Warfarin and share no interaction rows, so which one a clinician opens changes the warnings | RE-GRADED TO P1 AND CORRECTED 2026-08-17 on traced evidence, not inference. The original claim that the two records carry "three interaction rows each with ZERO in common" is FACTUALLY WRONG and understates the defect. Traced in data/medication-interaction-index.json: rows 0 (Pharmacokinetic, CYP2C9 inhibitors - amiodarone/fluconazole/metronidazole/cotrimoxazole) and 2 (Dietary, vitamin K) ARE substantively shared by both records. The real defect is COMPLEMENTARY INCOMPLETENESS, and it lands precisely on psychiatric drugs. warfarin-vka carries a CRITICAL CYP2C9 INDUCERS row (carbamazepine, St John's Wort - INR collapses toward 1.0, stroke risk) which warfarin-anticoagulant does NOT have. warfarin-anticoagulant carries a HIGH bleeding row for NSAIDs/Aspirin/SSRIs with twelve counterparties including six SSRIs (citalopram, escitalopram, fluoxetine, fluvoxamine, paroxetine, sertraline) which warfarin-vka does NOT have. So neither record is complete on its own, and each omits a psychiatrically central interaction the other holds. Resolution is strictly per-slug with no union anywhere: src/lib/medication-interactions.ts:207 and :241 read INDEX.bySlug[slug], and INDEX.names maps slug to display name, not name to slugs. Practical consequence for this app's actual user: a psychiatrist checking warfarin against sertraline sees the bleeding warning only if they happened to open warfarin-anticoagulant, and sees NOTHING if they opened warfarin-vka; checking warfarin against carbamazepine is the exact mirror. Nothing on screen distinguishes the two records - both display as "Warfarin". Secondary artefact worth fixing in the same pass: each record lists the OTHER as a counterparty (warfarin-vka's row-0 counterparties include warfarin-anticoagulant), so the duplicate is being treated as a drug that interacts with itself. Scope of this finding: it establishes the STRUCTURAL inconsistency only. Whether the interaction content of either record is clinically correct is not assessed here and is not an agent's call. Next: a clinician decides whether to merge the two records, delete one, or relabel them, and confirms the merged interaction set is complete - merging is the obvious candidate precisely because the two sets are complementary. Stop: do not patch the catalogue data automatically; this is clinical content. Grouping: treat alongside #318 as the clinical-content cluster - both are unvalidated medication-interaction surfaces in a prescribing tool. | PR #1923; docs/medication-interaction-lexicon-review.md flag section; tests/medication-interaction-lexicon-coverage.test.ts | 2026-08-13 | | #323 | P2 | task | 35 of 328 catalogue medications sit outside the resolved interaction graph, so the tool can never warn about them | Measured 2026-08-13 from data/medication-interaction-index.json using both endpoints of every row with a resolved counterparty: 35 of the catalogue's 328 medications sit outside the resolved interaction graph. They are concentrated in aperients (8), antibiotics (5), antidiabetics (4) and vitamins (3); psychiatry-relevant examples include topiramate and zolpidem. The former 127 count considered only inbound counterparty references and wrongly labelled source-only drugs such as celecoxib unreachable even though their own rows emit alerts. This is primarily CORPUS coverage: widening it requires authoring an interaction row or making existing source content machine-resolvable with clinical review, not indiscriminately widening lexicon selectors. PR #1923 closed the safety half - evaluateMedicationInteractions now reports unreachableCounterparties, composeMedicationVerdict treats it as incomplete so green is unreachable, and MedicationInteractionBlock names the uncovered drugs and says the absence of a warning is not evidence of safety. The generated list by class is the 'What this tool can never warn about' section of docs/medication-interaction-lexicon-review.md and refreshes with the report. Next: prioritise clinically relevant gaps on the prescribing surface. Stop: do not close this by loosening the matcher; that reintroduces the false-positive class (Sodium content, Vitamin K, hyperkalaemia prose) that was deliberately rejected. | PR #1923; docs/medication-interaction-lexicon-review.md coverage section; src/lib/medication-interactions.ts UNREACHABLE_SLUGS | 2026-08-13 | -| #324 | P1 | rec | No gate detects a merged PR whose content is silently reverted by a later merge resolution | **Outcome:** the file-level merge-loss detector is delivered; one authoritative row now tracks its remaining operational decision. **Delivered:** PR #1944 added scripts/audit-merge-loss.mjs through npm run audit:merge-loss and focused tests. It compares every changed file in a bounded main-history window with the landing commit's first parent, then reports possible reverts for human review. The implementation independently rediscovered the acf78bf casualties, including the #1803 token-retirement loss, and deliberately remains advisory because blob equality cannot distinguish a deliberate revert from an accidental merge-resolution loss. **Remaining:** decide whether it runs after merges or on a schedule, who triages positive findings, and whether the separate branch-versus-squash inbox-request-loss case should be a second detector or a mode of the same tool. A scheduled or required check without a named human disposition path would become ignorable noise. **Stop:** do not reimplement the delivered script, and do not make either detector blocking or auto-close findings until that ownership decision exists. | session 2026-08-13 blob sweep; PR #1944 audit implementation and tests; PR #1937 inbox-loss case; consolidated by PR #1956 review follow-up | 2026-08-13 | +| #324 | P1 | rec | No gate detects a merged PR whose content is silently reverted by a later merge resolution | **Outcome:** the file-level merge-loss detector is delivered; one authoritative row now tracks its remaining operational decision. **Delivered:** PR #1944 added scripts/audit-merge-loss.mjs through npm run audit:merge-loss and focused tests. It compares every changed file in a bounded main-history window with the landing commit's first parent, then reports possible reverts for human review. The implementation independently rediscovered the acf78bf casualties, including the #1803 token-retirement loss, and deliberately remains advisory because blob equality cannot distinguish a deliberate revert from an accidental merge-resolution loss. **Remaining:** decide whether it runs after merges or on a schedule, who triages positive findings, and whether the separate branch-versus-squash inbox-request-loss case should be a second detector or a mode of the same tool. A scheduled or required check without a named human disposition path would become ignorable noise. **Stop:** do not reimplement the delivered script, and do not make either detector blocking or auto-close findings until that ownership decision exists. SIGNAL-TO-NOISE CHARACTERISED AND TWO FIXES LANDED 2026-08-15 (owner-approved in session; still advisory, still unscheduled, still not blocking). (1) DEFECT FOUND AND FIXED: treeEntryReader split ls-tree output on the literal two-character sequence backslash-t rather than a tab, so the tree entry kept the filename. Same-path comparisons were unaffected, which is why the tool still found real losses, but isReconciliationMove compares an inbox path against its applied/ path, so the exemption could never match. Measured at 8069188 over 14 days: 51 findings / 255 flagged files / filesExempted 0, versus 11 findings / 66 flagged files / 189 exempted after a one-character fix - the exemption the script's own docstring says exists to stop inbox noise burying the genuine #1803 signal had been dead since it was written. Root cause of the escape: every test injected entryAt directly, so bare entries compared equal whether or not the path was stripped. Closed permanently by extracting parseTreeEntry as an exported pure function and testing it against real ls-tree output; mutation-verified (reintroducing backslash-t fails 3 tests plus the self-test). (2) MECHANISM CLASSIFIER ADDED: classifyRemoval walks the commits touching each flagged file between the landing and the ref, oldest first, takes the first whose tree entry already equals the pre-landing entry, and reports whether that commit was a merge (accidental) or single-parent (usually deliberate, and its subject says why). This is what makes the report triageable: over the window, 14 of 66 flagged files were merge-resolution removals with 13 from the single documented bad merge acf78bf, while all 52 others had explanatory single-parent subjects such as 'Re-land the --shadow-tight retirement', 'rework the viewer for phone and PWA reading' and 'ci: speed iteration without weakening gates'. Merge-resolution findings now sort first; unknown is reported rather than guessed. Mutation-verified in three directions (tab bug, newest-first walk, unknown-as-deliberate). GENUINE STILL-UNREPAIRED LOSSES, re-verified against main after it advanced past 8069188: #1800 fuzzy catalogue wiring is absent from therapies.ts, specifiers.ts and factsheets-data.ts AND all three of its tests carry zero fuzzy assertions so nothing can go red (tracked by #330); #1804's removal of UniversalSearchAlsoMatches from forms mode is reverted so the component is back at forms-search-results-page.tsx lines 44 and 894 with its guard assertions reverted, APPARENTLY UNTRACKED; #1796's ALLOWED_NODE_MAJOR_VERSIONS [24, 26] allowance is gone so worker/validate-runtime.ts still hard-codes nodeMajor() !== 24, APPARENTLY UNTRACKED; #1803 and #1807 lost design-system doc status rows while their code landed, so docs and code disagree. NEXT - the three decisions this row exists for are still open and are deliberately NOT implemented: (a) schedule, recommended weekly on a 14-day window rather than post-merge, because a post-merge trigger fires roughly 380 times per 14 days here and at merge time the loss has not happened yet; (b) triage owner, recommended routing to a pinned issue reusing the live-drift routing already covered by tests/live-drift-workflow.test.ts, with one named human, and not a required check; (c) recommended ONE tool with a --mode flag rather than a second detector, since the inbox case shares the window, landing enumeration and tree-entry comparison and differs only in paths and exemptions. Also recommended: the phantom-SHA class (a ledger record asserting a fix at 720e7027, an object that does not exist) is a DIFFERENT family - a ledger assertion with no landed content, checkable with git cat-file -e - and should get its own row rather than being folded into this tool. Stop unchanged: do not make either detector blocking or auto-close findings until the ownership decision exists. | session 2026-08-13 blob sweep; PR #1944 audit implementation and tests; PR #1937 inbox-loss case; consolidated by PR #1956 review follow-up | 2026-08-13 | | #325 | P3 | rec | A queued update request can silently clobber a row that changed after the request was written | **Outcome:** the inbox cannot apply a stale rewrite over someone else's newer content without anyone noticing. **Detail:** the inbox intake fixed ID allocation — ids are assigned at reconciliation, so two branches can no longer collide on a number, which was the sharper of the two hazards. It does not address content staleness. An 'update' request carries a full replacement '--detail' string written against whatever the author read at queue time; reconciliation applies it verbatim. If the target row changed on main between queueing and reconciling, the newer content is overwritten with no signal. The multiple-pending-mutations guard does not catch this: it fires only when two requests target the same id, not when one request is simply old. **Live near-miss, 2026-08-13:** a document-viewer ledger pass was drafted against a base four days stale, and its '#215' restatement was composed from that stale reading. It was caught only because the author re-read every row against current main before queueing — a discipline, not a gate. The same pass had already had to discard a directly-allocated '#295' because main had since claimed it; that half is now structurally impossible, this half is not. **Next:** consider fingerprinting the target row at queue time — the request schema is versioned ('version: 1'), so a 'baseRow' hash could be added to add/update/done payloads and compared at reconcile, refusing (or requiring an explicit override) when the row moved underneath. Weigh against just documenting the re-read discipline: this costs a schema bump plus writer, reconcile and self-test changes, and the failure needs a multi-day-stale base to bite. **Stop:** do not make reconciliation merge or three-way-diff detail text — a replacement that silently becomes a merge is harder to reason about than one that refuses. | session 2026-08-13 document-viewer ledger truth pass, PR #1930; scripts/ledger-inbox.mjs request schema | 2026-08-13 | | #326 | P3 | task | Keep post-restore environment recovery controls visible in the universal ledger | **Consolidated survivor for #188 and #196–#200 before their source rows are archived by PR #1920.** A schema restore is not operationally complete until all five environment-owned controls have been re-created and verified: (1) restore the ingestion, retention, and related `pg_cron` schedules and confirm they are active; (2) re-add required Supabase Vault secrets, including `cron_ingestion_jwt`, and verify names only without printing values; (3) re-set the required custom `app.*` database GUCs and verify them with read-only settings checks; (4) redeploy the required Supabase edge functions with the Deno v2.x toolchain and confirm the function list and health, only in an explicitly approved hosted-change window; and (5) re-enter dashboard-owned configuration, including auth providers and SSO redirect URLs, connection-pool caps, per-project keys, and `E2E_USER_*`, without committing secret values. **Next:** after every schema-restore drill or real restore, follow the disaster-recovery checklist in `docs/operator-backlog.md` and `docs/disaster-recovery-runbook.md`, record the verification outcome here, and keep the row open until all five controls are green. **Stop:** the runbooks are the execution procedure, not a substitute for this universal-ledger status row; do not treat a restored schema alone as recovered, expose secret values, or perform hosted writes without the required approval. | docs/operator-backlog.md disaster-recovery checklist; docs/disaster-recovery-runbook.md; #188/#196–#200; PR #1920 review | 2026-08-13 | | #327 | P3 | task | The recommended queue's Outcome cells are now unrendered dead text | **Residual of the queue-misdirection fix (PR #1902).** Both consumers — .claude/hooks/issues-surface.sh and scripts/issues-report.mjs — now derive each queue row's prose from the cited row's Detail cell, so the Outcome column reaches no reader through tooling. The stale prose still sits in the file, where a human opening it can read and act on it; for #231 that prose pointed at an approach the row had already recorded as refuted. **Implementation, established by building it 2026-08-13 — three findings that are not obvious:** (1) It cannot be a direct edit. check-ledger-write-discipline compares the canonical ledger against exactly applyRequestBatch(base, movedRequests), and no request type reaches the queue, so a hand edit is unlandable by construction. The rewrite has to live INSIDE applyRequestBatch — the function the checker itself imports — so checker and reconciler compute the same result; make it run for an empty batch and be idempotent so ordinary PRs are byte-identical. (2) It must land in the SAME commit as a reconcile. Code alone makes the checker compute normalise(base) while canonical stays un-normalised, failing every PR until a reconcile normalises it. (3) Do NOT drop the column, and do NOT blank composite rows. issues-report skips any queue row whose cells.length !== 7, so removing the column makes the queue vanish from /issues; and derivation deliberately skips composite ID(s) rows, so those still fall back to the Outcome cell and blanking it leaves them with no prose at all — filter to rows citing exactly one id. **Stop:** do not delete the queue table; order, acuity, capability, when and estimate exist nowhere else. | PR #1902; implementation attempt 2026-08-13 | 2026-08-13 | | #328 | P2 | issue | A row can outlive its own completion — nothing closes a ledger row when its work merges | **Found during the 2026-08-12 yield review; re-confirmed on main 2026-08-13.** The then-#304 row described a ranking-snapshot freshness fuse due to trip around 2026-08-19 and sat in the recommended queue as time-critical, but its work had already landed as commit d182844 (PR #1876) — the snapshot's generatedAt and sourceRunId no longer matched anything the row said. Nothing closes a row when its work merges: `issues:done` is a manual call, and the session that ships the work is often not the session that owns the row. This is the mirror of #292, which covers duplication BEFORE work starts; this is staleness AFTER it finishes, and it is more dangerous because the row keeps advertising urgency to every session that reads the queue. **Next:** the cheapest useful guard is a periodic re-verification pass that re-measures each open row against current main and flags rows whose stated evidence no longer reproduces — several rows already carry a hand-written VERIFIED CORRECT stamp, which shows the need but does it manually and unevenly. A stronger version has the handoff skill close the row in the same commit that lands the work. **Stop:** do not auto-close on keyword match; a row can be partially delivered (#215, #231) and auto-closing those would lose real remaining work. | session 2026-08-12 ledger yield review; re-verified 2026-08-13 | 2026-08-13 | | #329 | P2 | issue | All live mobile routes breach LCP; shared CSS delivery and JavaScript are the current bottleneck | PR #1927 is merged and deployed to Railway production at exact SHA f2abf5baf3f449a1803bedef9dc107f30b70db93. Three-sample live medians on that SHA are Documents 3374 ms, DSM 3961 ms, Forms 3507 ms, root 3819 ms, Therapy 3422 ms, and Services 3793 ms; desktop LCP is 580-679 ms and mobile CLS remains within the rule. The production CSS split is retained and reduced four canonical medians modestly, but every mobile route still breaches 2500 ms. Root trace attribution is now concrete: TTFB 283 ms, LCP render delay 3449 ms, the 46,724-byte transferred shared stylesheet completes at 3644 ms under the throttled critical-request contention, total main-thread work is 1785 ms, script evaluation is 1030 ms, and shared chunk 8322 alone consumes 870 ms CPU. This is separate from canonical #117, which continues to track the unresolved Therapy catalogue payload and per-field safety decision. Next: split the 4,251-line global stylesheet by route ownership and reduce the shared search-shell/root client boundary before repeating the same bounded live matrix. Therapy field safety review remains required for search/pathways. INP remains unverified because Lighthouse does not measure it and no usable CrUX result exists. Stop: do not strip clinical fields, weaken the Lighthouse budget, refresh a passing baseline to hide latency, or claim an INP pass. | PR #1927; Railway deployments 1224ed55-210d-443b-94e5-20f87475468c and 810cc8b3-e39a-493f-b18f-8c63d150d53f; live Web Vitals runs 31719448766 and 31719451951; PR #1933 review | 2026-08-13 | -| #330 | P2 | task | Re-land PR #1800 (fuzzy catalogue search), applying the #310 one-edit cap in the same commit | PR #1800 squash-merged as 022c83b on 2026-08-10 and its entire content is absent from main: git show origin/main:src/lib/catalog-search.ts \| grep -c typoDistanceLimit returns 0, eight of its 11 source and test files are byte-identical to their pre-#1800 state. The remaining three (`src/components/therapy-compass/data/select.ts`, `src/lib/formulation.ts`, and `tests/formulation.test.ts`) contain later unrelated changes, but the fuzzy-search hunks are absent from them too; preserve those newer changes during the re-land. Cause and evidence in the merge-loss detector row filed alongside this one. Consequence today is a MISSING FEATURE, not a live hazard: because the matcher is gone, the #310 cross-drug defect is not reachable on main. Do not close #310 on that basis, and do not re-land #1800 unchanged. RE-LAND WITH THE FIX: #310 measured that the tier term.length >= 8 -> 2 edits is the problem, because Damerau scores an adjacent transposition as one edit, so fluoxetine to duloxetine is distance 2 and both are ten characters. Re-run 2026-08-13 against the algorithm confirms it, and confirms prednisone to prednisolone as the second real cross-drug hit. Capping that tier at 1 edit removes both while preserving sertraline to sertralin style recovery. The row's other claims also held on re-run: citalopram and escitalopram do not fuzzy-match, because the substring guard fires first, and clozapine/clonazepam and quetiapine/olanzapine are correctly out of range. Next: cherry-pick 022c83b onto current main, change typoDistanceLimit's >= 8 tier from 2 to 1, and add a test over real catalogue drug names with both the exact and the near-match record present, asserting the wrong drug is excluded while the exact drug remains. Gate: focused Vitest on `tests/catalog-search.test.ts` plus the other four test files #1800 touched. Stop: this path is clinicalRisk true under classifyPullRequestFiles because catalog-search.ts feeds medications.ts and prescribing, so the PR needs a complete Clinical Governance Preflight and must not be bundled with unrelated chores. ragRanking is correctly false; this is catalogue ranking, not pgvector retrieval. | session 2026-08-13; 022c83b; origin/main at 63526ee; row #310; algorithm re-run locally against real drug-name pairs | 2026-08-13 | -| #331 | P2 | issue | check:medication-lexicon-report fails on 3 independent branches despite zero diff on the flagged file or its inputs | **Outcome:** one authoritative owner for the medication-report staleness problem, including its missing CI coverage. **Evidence:** on 2026-08-14, check:medication-lexicon-report reported the review document stale on three independently authored branches (#1947, #1949, #1950) although the document, lexicon sources, medication snapshot, and interaction index were untouched. The symptom must therefore be investigated against a clean current main rather than fixed opportunistically in unrelated work. **Scope:** this row also carries the CI evidence formerly duplicated in #333: the check is reached only at the end of verify:pr-local and no workflow invokes it, so CI can stay green while a local PR preflight fails. **Next:** inspect the generator and its staleness comparison against current main; if the report is genuinely stale, regenerate it in a dedicated clinical-document change, otherwise fix the comparison. In the same decision, either make the validated check part of the appropriate CI contract or move it out of the local preflight so its enforcement matches its ownership. **Stop:** do not delete or weaken the check merely to green an unrelated preflight, and do not regenerate a clinical-facing artifact without checking whether the diff changes clinical content. | PR #1947, PR #1949 and PR #1950 clean-branch reproductions; PR #1942 preflight; session 2026-08-14; consolidated by PR #1956 review follow-up | 2026-08-14 | | #332 | P3 | task | Three mode-nav icon glyphs sit at 17px, off the --spacing-icon-* scale, and no gate flags them | Split out of #275 rather than folded into its badge-box token. mode-nav/mode-nav.tsx:64 and :214 and mode-nav/nav-slot-ink.tsx:44 size their with h-[1.0625rem] w-[1.0625rem] — 17px against an icon scale of 12/14/16/20/24 (--spacing-icon-xs..xl in the globals.css @theme block). #275 counted these among its five files because they share the badge's number, but they are a different role: the badge is a text-bearing box sized around its own --text-2xs numeral, these are glyphs. They are now the only consumers of that value, since the badge moved to --spacing-search-band-badge. Nothing gates this: check-icon-scale.mjs enforces only the retired 4.5 (18px) half-step and its header states it deliberately does NOT flag arbitrary h-[Nrem], because non-icon boxes legitimately use that form. So this is unguarded and will not self-report. Why it was not just fixed: snapping to size-icon-md (16px) or size-icon-lg (20px) visibly changes nav chrome at every breakpoint, and 17px is close enough to 16 that the choice looks arbitrary without seeing it rendered — a design call, not a token swap. Next: get a Chromium look at mode-nav at phone and desktop widths with the icon at 16 and at 20, pick one, then migrate all three together. If 17px turns out to be deliberate, say so in a comment at the call site and consider whether check:icon-scale should flag off-scale arbitrary icon sizes on -typed elements specifically, which would have surfaced this. Stop: do not add a 17px step to --spacing-icon-* to make the problem go away — that token block's own comment argues against widening the scale off the 4px grid, and it would sanction the drift rather than resolve it. | session 2026-08-14; split from #275; check-icon-scale.mjs header | 2026-08-14 | | #334 | P3 | issue | Claude Code web containers can ship Node 22 with no node_modules, so npm ci fails engine-strict before any work starts | Hit 2026-08-14 at the start of a Claude Code on the web session, and it blocks a session completely until worked around, so it is worth recording even though the cause is the container image rather than this repo. The container provided /opt/node20, /opt/node21 and /opt/node22 with node22 on PATH, no nvm, and no node_modules in either the primary checkout or a fresh worktree. package.json requires node >=24.15.0 <25 with engine-strict, so 'npm ci --include=dev' aborts immediately with 'notsup Required: {node: >=24.15.0 <25, npm: 11.x} Actual: {npm: 10.9.7, node: v22.22.2}'. Nothing in the repo can fix this from inside, because the failure happens before any repo script can run — .nvmrc correctly says 24 and is simply not consulted, and there is no nvm for it to drive. Workaround used, which took about a minute and is safe: fetch the current 24.x from the nodejs.org dist index, untar to /opt/node24, and prefix subsequent commands with 'export PATH=/opt/node24/bin:/opt/node24/bin:/root/.local/bin:/root/.cargo/bin:/usr/local/go/bin:/opt/node22/bin:/opt/maven/bin:/opt/gradle/bin:/opt/rbenv/bin:/root/.bun/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin'. Everything downstream then behaved normally — npm ci, the full unit suite, build, and the Playwright-free gates all passed. Worth knowing that this is a DIFFERENT surface from the Codex Cloud provisioning path: scripts/setup-codex-cloud.sh and scripts/setup-codex-worktree.mjs cover Codex, and docs/codex-cloud.md is explicit that Cloud mirrors the tracked toolchain, but neither runs for a Claude Code web session, so that hardening does not carry over. Next: decide whether this deserves repo-side help at all. Options are a short note in the AGENTS.md or CLAUDE.md orientation telling an agent to install Node 24 to /opt/node24 and re-export PATH rather than concluding the environment is broken, or a small bootstrap script equivalent to the Codex ones that a web session can run first. Prefer the note: a bootstrap script that downloads a runtime is a bigger surface than the problem. Stop: do not relax the engines range, drop engine-strict, or pass --force to get npm ci through — the Node 24 floor is enforced deliberately in several places (preinstall, check:runtime, scripts/dev-free-port.mjs) and loosening it to accommodate a bad container would disable a real guard. | session 2026-08-14; Claude Code web container for PR #1942 | 2026-08-14 | | #336 | P3 | rec | Decide whether responsive breakpoint windows get named tokens, or stay raw min-[]/max-[] everywhere | Split out of #275 rather than guessed at. The repo defines ZERO --breakpoint-* tokens, and at least nine sites hand-write the arbitrary form: min-[414px]:max-[429px] at clinical-dashboard/result-filter-control.tsx:231, plus max-[359px] (search-heading-mockups, differentials/diagnosis-map-panel.tsx:1036, clinical-dashboard/account-setup-dialog.tsx:98) and max-[389px] (factsheets/factsheets-search-page.tsx:176, clinical-dashboard/search-results-header-band.tsx:532, factsheets-compact-view-mockups). #275 asked for the 414-429 window to be tokenised alongside the badge box; that was deliberately NOT done, because naming one window while eight peers stay raw reintroduces exactly the one-call-site drift #275 exists to stop, just on a different axis. This is a real decision with two defensible answers and it should be made once, for all of them. (a) Stay raw and say so in docs/design-system/GATES.md: the values are per-device band edges carrying measured justifications in their own comments, they are not a scale, and a Tailwind 4 --breakpoint-* entry adds BOTH the min and max variant to every utility in the build for a single consumer. (b) Name them: Tailwind 4 --breakpoint- generates : and max-:, so the 414-429 window needs two entries (414px and 430px, since max-[429px] is inclusive and max- is exclusive), and 359/389 would want their own. Note the mockup hits are design scratch and out of scope for any gate. Next: pick (a) or (b), record it in GATES.md section 3 so the next session does not re-derive it, and only then migrate. Stop: do not migrate one window ahead of the decision. | session 2026-08-14; split from #275 during the design-token relands PR | 2026-08-14 | @@ -221,6 +209,13 @@ removed after current-main verification; it is not missing recommended work. | #341 | P2 | task | Route the remaining ~22 unguarded source-slice test windows through the guarded helper | PARTIALLY DONE 2026-08-15 by PR #1985, which added tests/helpers/source-contract.ts and migrated the three worst files. The hazard this closes is a silent pass, not fragility: the idiom source.slice(source.indexOf(start), source.indexOf(end)) returns -1 for a missing end marker, and slice(n, -1) does not throw — it returns the rest of the file bar one character. A renamed end marker therefore converts a scoped assertion into a whole-file assertion and every positive toContain in it keeps passing for the wrong reason. The mirror case, a missing start, yields slice(-1, n) so every negative assertion passes vacuously. Neither shows up as a failure. The helper throws on a missing start, a missing end, and an AMBIGUOUS start (a window anchored on a string that appears twice silently covers only the first hit). MIGRATED: search-route-ownership.test.ts (three windows, including one whose end marker is an indentation depth and one anchored on a comment string), document-detail-performance.test.ts (end marker was the next literal 'useEffect' token, of which DocumentViewer has several), therapy-compass-responsive-contract.test.ts. REMAINING, roughly 22 windows across ~11 files, none migrated: audit-navigation-auth-regressions.test.ts is the densest and was deliberately skipped because PR #1983 edits the same file and the anti-churn rule prefers one late sync to a merge fight — do it once #1983 lands. Also tools-search-directions-mockups.test.ts, in-page-nav-playwright-contract.test.ts and document-section-nav-contract.test.ts (the last two slice ui-smoke.spec.ts between Playwright test titles, so renaming OR reordering an unrelated spec silently rescopes them), and rag-retrieval-parallelism.test.ts, which is left for a session that flags the RAG surface first per AGENTS.md. A REAL COVERAGE HOLE was found while surveying and is NOT yet fixed: audit-navigation-auth-regressions.test.ts around line 285 anchors on '{showUniversalAlsoMatches &&', which occurs TWICE in ClinicalDashboard.tsx (the second around line 3832 sits outside the window), so its not.toContain check does not enforce the named contract across the file. The new helper would reject that anchor outright, which is how it was found. Fix it in the same pass as that file's migration. STOP: do not loosen the demo-data boundary pins in favourites-demo-boundary.test.ts. The exact conditional-spread form '...(demoMode ? prototypeFavouriteItems : [])' with its paired negative is the live-vs-demo privacy contract, and its strictness is the point. Likewise leave the SQL windows ending on '$$;' and header-scroll-hide-contract.test.ts anchoring on the matching '' closing tag — those are true structural terminators and are the model the rest should move toward. | session 2026-08-15; PR #1985; tests/helpers/source-contract.ts | 2026-08-15 | | #342 | P2 | issue | Recurring 'Unhandled server request error' on /api/search and /api/search/universal is untriaged | Three Sentry issue groups in clinibase-xz over 24h (JAVASCRIPT-NEXTJS-Y, -Z, -10), 17 events, 0 users impacted, all titled 'Error: Unhandled server request error' with culprit chunk 1261.js:2:4801. Top frames are /api/search/route.js and /api/search/universal/route.js. First seen 2026-08-14T08:44:37Z on release c9b089c92c975297c10649b005401d5ae337cf48, roughly six hours BEFORE PR #1946 merged, so it is not caused by the retrieval row contract; the post-merge group is the same error refingerprinted by the release change. The error string does not appear anywhere in repo source, so it likely originates in a dependency or an instrumentation wrapper — origin unidentified. Nobody owns this. Next step: identify what throws it, then decide whether it is a bot/scanner artefact or a real request-handling gap. | Sentry clinibase-xz, reviewed 2026-08-15 | 2026-08-15 | | #343 | P3 | task | Make the retrieval row contract's source_metadata pin structural, not data-guaranteed | rag-row-contracts.ts pins source_metadata to a JSON object via z.record(...), but documents.metadata is bare jsonb and permits arrays and scalars. Measured against the live project (sjrfecxgysukkwxsowpy) on 2026-08-15: all 2851 documents are object-typed, so nothing breaks today and no live errors exist. The guarantee is data, not schema — a future ingest path could violate it and take retrieval down for that document's chunks. Fix is either a check (jsonb_typeof(metadata) = 'object') constraint on public.documents, or loosening the pin. Every other required field in that contract is backed by a not-null constraint. | PR #1946 review + live Supabase verification 2026-08-15 | 2026-08-15 | +| #6BG9X2 | P2 | task | R2 + R3: claim-support strictness rejects verbatim-faithful guideline restatements (directive normativity; topic-overlap dilution) — packet S1c | R2: normativeDirectiveActions in src/lib/rag/rag-claim-support.ts has no pattern for 'usual / recommended ... dose is ...' guideline phrasing, so imperative claims ('start lithium at 500 mg nocte') fail against descriptive norms; reproduced offline on the EMHS lithium chunk. R3: a claim synthesising two adjacent source bullets fails the >=50% single-segment topic-overlap requirement even when every atom matches; reproduced offline. Fix R2 with a small pattern addition plus adversarial negatives; MEASURE R3 before loosening. One PR after S1b merges and its canary is green; RAG impact behaviour change; canary pair. Packet: docs/rag-improvement/HANDOVER.md S1c. Stop: no grounding-gate weakening beyond the two named artefacts. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | +| #0MSNT8 | P3 | task | Governance Option B decided: tag document-summary rows with similarity_origin 'document_context', keep the confidence label — implement as packet G1 | Owner decision 2026-08-17 on the question queued by #212 tranche 3 (buildDocumentSummaryResults stamps similarity: 1 with no similarity_origin; deriveConfidence excludes only synthetic_text). Option B: add a new similarity_origin value document_context to the union in src/lib/types.ts and to src/lib/answer-stream-contract.ts, stamp it in buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts), keep deriveConfidence unchanged and pin it with discriminating tests, keep rag.ts synthetic_similarity_count from counting it, and update docs/clinical-hazard-analysis.md H5a. Rationale: the only caller is the document-summary route where the query is the document itself and citation support is still verified; Option A (tag as synthetic_text so summaries cap at medium) rejected as a label downgrade without a measured safety gain. RAG impact: no retrieval behaviour change; no canary. Packet: docs/rag-improvement/HANDOVER.md G1. Closes the P1 governance question row once landed. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | +| #DP6M3G | P1 | task | R1: unbudgeted strong escalation makes provider_timeout the dominant lithium fallback — route the dosing class to strong before the deadline (packet S1b) | S1 (PR #2022) post-fix live probes: 'Lithium dosing?' 4/4 source-only, 3/4 as provider_timeout. fast_unsupported_retry_strong launches a strong generation into the fast route's leftover ~10-13 s; only the truncation self-heal is deadlineAllowsGenerationRetry-gated. Ladder rung 3 (README A1): route medication_dose_risk / dosing to the strong route in chooseAnswerRoute (src/lib/rag/rag-routing.ts) BEFORE the route deadline is created — not in shouldRetryWithStrongAfterFast, and NOT a budget change (#231 stop condition stands). Own PR, RAG impact behaviour change, canary pair, Clinical Governance Preflight, check:production-readiness. Owner decided 2026-08-17 this lands before S2 (A2/A3 add length; length under the unbudgeted retry pushes more dosing queries into timeout). Packet: docs/rag-improvement/HANDOVER.md S1b. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | +| #BTVMVK | P2 | issue | Recurring 'Unhandled server request error' on /api/search and /api/search/universal in Sentry — unowned, pre-dates #1946 | Sentry (clinibase-xz): three issue groups in 24h, 17 events, 0 users impacted, on /api/search and /api/search/universal, all with culprit chunk 1261.js:2:4801. First seen 2026-08-14T08:44:37Z on release c9b089c9, about six hours before PR #1946 merged, so not caused by the row contracts. Recorded in the #212 tranche-1 handover; never captured durably until now. Next: triage the Sentry groups (read-only Sentry MCP or dashboard), map chunk 1261.js to source via the release's source maps, reproduce locally with the request shapes Sentry recorded. Stop: do not silence the error path; search routes are clinical output. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | +| #J912J9 | P1 | issue | Decide whether a fabricated similarity of 1 on document-summary rows may earn the high confidence label a clinician reads | buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts) stamps similarity: 1 on document-summary rows -- a fabricated score, not a measured cosine -- and does NOT set similarity_origin. deriveConfidence (src/lib/rag/rag-answer-support.ts:32-36) computes strongestNonSynthetic by EXCLUDING rows tagged similarity_origin === "synthetic_text", so an untagged fabricated 1.0 IS counted, and line 36 gates the "high" verdict on strongestNonSynthetic >= 0.82 with at least 2 accepted citations. A document summary therefore reaches the "high" confidence label on a score nobody measured. Note the mechanism precisely: buildDocumentSummaryResults does NOT set the tag -- the three call sites that DO tag synthetic scores are in rag-candidate-sources.ts (lines 626, 721, 934), and it is the ABSENCE of the tag here that admits the fabricated score. docs/clinical-hazard-analysis.md H5a records the adjacent document-lookup fast-path hazard. QUESTION: should a fabricated similarity count toward "high"? Adding the tag would demote these answers to medium or low. This is a clinical-governance decision before it is an engineering one: either way it needs its own design, discriminating offline tests that separate tagged from untagged rows, and a live eval-canary before/after pair per docs/rag-behaviour/. Deliberately excluded from PR #1981 and from the #212 tranche 3 PR rather than bundled, because it changes clinical output. Verified on main at d0276718. | ledger #212 tranche 3 (src/app/api row contracts), session 2026-08-17 | 2026-08-17 | +| #TYJ0XP | P3 | rec | eval-canary.yml is post-merge only (repository_dispatch + Sunday cron, no ref input) — record this in docs/rag-behaviour so sessions stop expecting a branch canary | .github/workflows/eval-canary.yml triggers on repository_dispatch type eval-canary and schedule cron 0 18 * * 0; it always loads the default branch and has no workflow_dispatch or ref input, so a canary can only ever measure main. Consequence for the RAG programme: 'canary pair' means latest green run on main before the merge -> a dispatch after the merge (gh api repos/BigSimmo/Database/dispatches -f event_type=eval-canary), compared with npm run eval:retrieval:compare -- --fail-on-regression on the eval-canary-output artifacts. Recorded correctly for S1 (baseline run 31964560921 -> post run 32025082010, zero regressions). Next: add one paragraph to docs/rag-behaviour/safeguards.md canary protocol; no workflow change. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | +| #ND10QT | P3 | rec | source_metadata pin in rag-row-contracts.ts is data-backed only — add check (jsonb_typeof(metadata) = 'object') on documents or loosen the pin | src/lib/rag/rag-row-contracts.ts requires source_metadata to be a JSON object (z.record) while documents.metadata jsonb permits arrays and scalars. Live query 2026-08-14 (project sjrfecxgysukkwxsowpy): select jsonb_typeof(metadata), count(*) from public.documents group by 1 returned a single row, object = 2851, so nothing breaks today — but the guarantee is data, not schema, and a future ingest could make retrieval throw RetrievalRowShapeError for that document. Options: a check constraint via a new migration (role postgres; run check:migration-role) or loosen the pin. The module's doc comment also claims every required field is 'not null in supabase/schema.sql', which is true for nine fields and false for source_metadata — one-clause docs fix in the same PR. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | ## Resolved / archive @@ -479,3 +474,10 @@ Move resolved rows here with the resolution date and a one-line outcome. Keep th | #163 | task | Redesign Services search results (Progressive Referral Workflow) | CLOSED 2026-08-15 by PR #1982, but the row was half-stale when picked up and the split is worth recording. ALREADY ON MAIN before this PR, verified against e60b49a: the query is the

via SearchResultsHeaderBand headingLevel={1} (the count is separate neutral text, never a heading), the shortlist bar is already conditional on selected.length, Compare already exists, and the four-card numbered walkthrough (ServiceReferralFlow) is mounted only on /services/[slug] — tests/ui-tools.spec.ts:1394 already asserted it absent from results. The current/ baseline PNG the row was written against is from 2026-07-31 and no longer reflects main, which is why the row read as untouched work. DELIVERED BY THIS PR: the tiny Search/Shortlist/Compare/Refer dot rail (new ServiceReferralProgress, accessible name Referral progress — deliberately not Referral workflow, so the absence assertion cannot be satisfied by renaming the old component back onto the route); row compaction dropping the Catchment/Eligibility/Cost strip and the confidence pill (both values remain untruncated on the record behind Review referral); a per-row bookmark wired to account favourites and kept distinct from the non-persisted shortlist, with a visible polite status because sign-in-required is the common guest outcome; and the shortlist banner moved below the heading it qualifies. NOT DONE, deliberately: the comp draws the query dominant and the count small, which is the reverse of the shipped band. That weighting is a documented contract (docs/search-chrome-behaviour.md Results band rules 1-2), shared by twelve modes, and visual-baselined from /services?q=CMHT&run=1, so inverting it is a repo-wide change and not this row. The outcome this row asked for holds either way. Captured separately. Verification: verify:pr-local all ten selected gates passed; UI proof delegated to CI Production UI because the container ships chromium-1194 against a pinned 1234 (#255/#312). | 2026-08-15 | | #164 | task | Redesign Favourites as hybrid dashboard + search (no ModeHome) | CLOSED 2026-08-15 by PR #1983. /favourites is one dashboard+search workspace; no ModeHome was reintroduced and no separate Favourites home route exists. DELIVERED: the 'Favourites command library' marketing H1, its heart icon tile and its explanatory subtitle are retired for a plain Favourites heading with the item count beside it as non-heading text; the desktop FavouritesSidebar plus the phone FavouritesMobileQuickViews and FavouritesMobileBrowseRail collapse into one chip rail carrying sets, Pinned/Source-backed and types with counts, after which favourites-library-nav.tsx had no callers and was deleted (529 lines, and it paid down one gate-8 edge conflict as a side effect: 25 to 24); the empty query shows Continue then Recent and Your sets side by side, a typed query demotes that band to a collapsed disclosure and filters the table in place with an 'N matches for ...' header; and the redundant filter computation is gone — filteredItems and the table's own tableRows were derived independently from identical inputs, so the band count and the table count were two answers to one question, and the page now derives rows once and passes them down. IN-PLACE FILTERING IS REAL, not just same-surface: the page reads the shared composer's live draft via useSearchCommand seeded from the route's submitted ?q=, the same pattern tools-search-results-page already uses, so typing filters without navigating and without a second input and the one-composer contract is untouched. NOT DONE, deliberately and worth knowing: Continue and Recent still render from the existing derivation, where lastUsedByItemId and pinnedItemIds are hard-coded five-entry literals for demo slugs and real registry items fall back to the literal string 'Saved'. Direction B leads with both surfaces, so they want a genuine per-item last-opened timestamp; that is a data-layer change and building it here would have silently rescoped this row. Captured as its own row. Verification: verify:pr-local all ten selected gates passed, unit suite 607 files / 6584 passed; UI proof delegated to CI Production UI (chromium-1194 vs pinned 1234, #255/#312). | 2026-08-15 | | #293 | issue | Gate 2 needs a phone-viewport deterministic surface; the `min-h-tap` 0px finding is REFUTED | Closed after PR #1962 landed on main (commit 885c613): tests/ui-style-contract.spec.ts now has a deterministic Gate 2 tap-carrier enumeration test at a 390x844 phone viewport on /forms's static home, polling until 3 consecutive reads agree. Finding 1 (min-height override) was already refuted as intentional (sm: desktop release of the phone-only floor); this closes Finding 2, the last open part. | 2026-08-15 | +| #212 | task | Replace as unknown as casts and unvalidated JSON.parse with Zod or runtime guards | Closed 2026-08-17 by tranche 4 PR #2037 (worker/main.ts), completing the four-tranche sweep: T1 PR #1946 (src/lib/rag/rag.ts via rag-row-contracts.ts, live-verified), T2 PR #1981 (rag-candidate-sources.ts), T3 PR #2023 (src/app/api/** via src/lib/validation/row-contracts.ts, squash 440a34f71), T4 PR #2037 (worker/row-contracts.ts). Corrected population baseline from the T3 audit (folding in its queued row correction, superseded request 2f5fbcde): measured on main at d0276718 across src/, worker/ and scripts/ there were 80 'as unknown as' and 60 JSON.parse occurrences, but most are legitimate outbound serialization (pgvector query_embedding args, telemetry inserts cast to Json) or client handles, so the remediable number was far lower; src/lib/rag retains 5 deliberate outbound/client-handle casts; T3's audit of all 41 API route files found only 3 'as unknown as' (outbound telemetry, correctly left) and zero JSON.parse, with every request body already Zod-validated via parseJsonBody - the 4 genuine inbound targets (search_document_chunks RPC rows, document_chunks table fallback, document_labels rows, search_schema_health payload) got row contracts. T4 audit of the final 12-cast worker population: 1 inbound cast (claim_ingestion_jobs rows incl. documents to_jsonb payload) replaced with per-row fail-soft validation that terminally fails malformed rows through fail_or_retry_ingestion_job instead of throwing outside the job lifecycle; 2 deep-memory parameter casts replaced by asserting loadEnrichmentRows read-backs (contained by the inline-enrichment try, record-and-skip); 9 outbound insert/Json/CJS-interop casts deliberately left with per-site reasons in the PR. Scripts are not production paths. Residual, distinct population outside this item's inventory: unguarded bare-as casts on extractor tableRows (worker/main.ts:693/1004) and OpenAI vision classification widening (worker/main.ts:798-800/1100/1115) - raise separately if they warrant their own item. | 2026-08-17 | +| #237 | rec | Eyeball low-confidence AccessibleTable densities at 320px before freezing Linux visual baselines | Completed with a real Chromium mockup journey at 320x700 using the production AccessibleTable and MissingValue components. The focused repository wrapper passed twice; computed proof showed normal whitespace, no ellipsis, both full 'Not recorded' phrases within the table wrapper, wrapper scroll delta at most 1px, and zero page overflow. The final 320x700 screenshot was visually inspected and confirmed the low-confidence warning and phrases are legible and uncut. No product CSS adjustment was needed. | 2026-08-15 | +| #330 | task | Re-land PR #1800 (fuzzy catalogue search), applying the #310 one-edit cap in the same commit | Resolved on current main 512b8c20e by commit 247a35936cfdc78e2e85dd889d5d9dbdc3804e32: conservative fuzzy catalogue search and the one-edit cap are implemented with wrong-drug regression coverage. | 2026-08-15 | +| #331 | issue | check:medication-lexicon-report fails on 3 independent branches despite zero diff on the flagged file or its inputs | Resolved on current main: PR #1951 commit 9628b246 repairs the generated medication lexicon report and PR #1979 commit 17402395 wires check:medication-lexicon-report into the static_heavy_changed CI contract. Clean current main 512b8c20 reproduces the check green at 28 catalogue terms; clinician review remains separately tracked by #318. | 2026-08-15 | +| #192 | task | X6: Raise clinical/retrieval/answer coverage floors | Resolved by PR #1964 (commit adc5182a7): X6 added four clinical/retrieval/answer coverage groups plus global/broad floors and CI enforcement, raising evidence/verification branches 81 to 84 and core RAG branches 72 to 76 without lowering any floor. Current-main static CI scope check passes. A fresh exclusive coverage run is useful after later RAG commits but is freshness evidence, not remaining X6 implementation. | 2026-08-15 | +| #162 | task | Redesign Tools search results state (Compact Results Instrument) | Resolved on current main 512b8c20e by commit 88e3117ff3e078f0c88cd2e5e5c360ce4efca665: Tools is an all-results directory with query-as-H1, shared results band, dense rows, filters, and one-composer ownership. Stale local branch disposition remains separate hygiene work. | 2026-08-15 | +| #238 | rec | Visual pass for Sheet portal default on settings, sidebar, and answer overlays | Completed on clean current main 512b8c20 with real Chromium journeys across all five required Sheet host contexts. Repository wrapper result: 5 passed in 23.8s, covering settings via ClinicalSidebar at phone widths with close/Escape, source-backed answer Sources/safety sheets with focus return, answer Sources/Clinical notes/Evidence sheets, the mobile launcher detail sheet, and the form section-navigation sheet. This is host-level browser evidence, not generic Sheet unit coverage. | 2026-08-15 |