diff --git a/docs/outstanding-issues-inbox/0177f340-3275-472f-8ef3-3a415e335125.json b/docs/outstanding-issues-inbox/applied/0177f340-3275-472f-8ef3-3a415e335125.json similarity index 100% rename from docs/outstanding-issues-inbox/0177f340-3275-472f-8ef3-3a415e335125.json rename to docs/outstanding-issues-inbox/applied/0177f340-3275-472f-8ef3-3a415e335125.json diff --git a/docs/outstanding-issues-inbox/01cacbb4-9d47-4d05-bf34-becd134b9075.json b/docs/outstanding-issues-inbox/applied/01cacbb4-9d47-4d05-bf34-becd134b9075.json similarity index 100% rename from docs/outstanding-issues-inbox/01cacbb4-9d47-4d05-bf34-becd134b9075.json rename to docs/outstanding-issues-inbox/applied/01cacbb4-9d47-4d05-bf34-becd134b9075.json diff --git a/docs/outstanding-issues-inbox/02846e88-de52-42be-8c67-a5335e722fea.json b/docs/outstanding-issues-inbox/applied/02846e88-de52-42be-8c67-a5335e722fea.json similarity index 100% rename from docs/outstanding-issues-inbox/02846e88-de52-42be-8c67-a5335e722fea.json rename to docs/outstanding-issues-inbox/applied/02846e88-de52-42be-8c67-a5335e722fea.json diff --git a/docs/outstanding-issues-inbox/39b3c784-e840-4308-91c2-a146f4ff5c45.json b/docs/outstanding-issues-inbox/applied/39b3c784-e840-4308-91c2-a146f4ff5c45.json similarity index 100% rename from docs/outstanding-issues-inbox/39b3c784-e840-4308-91c2-a146f4ff5c45.json rename to docs/outstanding-issues-inbox/applied/39b3c784-e840-4308-91c2-a146f4ff5c45.json diff --git a/docs/outstanding-issues-inbox/3bb892c3-a2c5-4718-828a-1ccedface10e.json b/docs/outstanding-issues-inbox/applied/3bb892c3-a2c5-4718-828a-1ccedface10e.json similarity index 100% rename from docs/outstanding-issues-inbox/3bb892c3-a2c5-4718-828a-1ccedface10e.json rename to docs/outstanding-issues-inbox/applied/3bb892c3-a2c5-4718-828a-1ccedface10e.json diff --git a/docs/outstanding-issues-inbox/3d23b389-3d2b-4043-b3b1-85633b5962a1.json b/docs/outstanding-issues-inbox/applied/3d23b389-3d2b-4043-b3b1-85633b5962a1.json similarity index 100% rename from docs/outstanding-issues-inbox/3d23b389-3d2b-4043-b3b1-85633b5962a1.json rename to docs/outstanding-issues-inbox/applied/3d23b389-3d2b-4043-b3b1-85633b5962a1.json diff --git a/docs/outstanding-issues-inbox/43f5e351-6a7b-4217-b4a4-897d8be9b905.json b/docs/outstanding-issues-inbox/applied/43f5e351-6a7b-4217-b4a4-897d8be9b905.json similarity index 100% rename from docs/outstanding-issues-inbox/43f5e351-6a7b-4217-b4a4-897d8be9b905.json rename to docs/outstanding-issues-inbox/applied/43f5e351-6a7b-4217-b4a4-897d8be9b905.json diff --git a/docs/outstanding-issues-inbox/4aa68f9f-0961-4681-9859-dd0560a1bbbf.json b/docs/outstanding-issues-inbox/applied/4aa68f9f-0961-4681-9859-dd0560a1bbbf.json similarity index 100% rename from docs/outstanding-issues-inbox/4aa68f9f-0961-4681-9859-dd0560a1bbbf.json rename to docs/outstanding-issues-inbox/applied/4aa68f9f-0961-4681-9859-dd0560a1bbbf.json diff --git a/docs/outstanding-issues-inbox/53491e9a-64f0-4e67-b175-1e69c82a1be4.json b/docs/outstanding-issues-inbox/applied/53491e9a-64f0-4e67-b175-1e69c82a1be4.json similarity index 100% rename from docs/outstanding-issues-inbox/53491e9a-64f0-4e67-b175-1e69c82a1be4.json rename to docs/outstanding-issues-inbox/applied/53491e9a-64f0-4e67-b175-1e69c82a1be4.json diff --git a/docs/outstanding-issues-inbox/5b3c1df6-69e3-4f90-877f-915ee75358de.json b/docs/outstanding-issues-inbox/applied/5b3c1df6-69e3-4f90-877f-915ee75358de.json similarity index 100% rename from docs/outstanding-issues-inbox/5b3c1df6-69e3-4f90-877f-915ee75358de.json rename to docs/outstanding-issues-inbox/applied/5b3c1df6-69e3-4f90-877f-915ee75358de.json diff --git a/docs/outstanding-issues-inbox/5b6fdafd-91f9-4225-a472-52b4a5d16303.json b/docs/outstanding-issues-inbox/applied/5b6fdafd-91f9-4225-a472-52b4a5d16303.json similarity index 100% rename from docs/outstanding-issues-inbox/5b6fdafd-91f9-4225-a472-52b4a5d16303.json rename to docs/outstanding-issues-inbox/applied/5b6fdafd-91f9-4225-a472-52b4a5d16303.json diff --git a/docs/outstanding-issues-inbox/65a686fe-7165-4fc9-9f1b-4b98c1824203.json b/docs/outstanding-issues-inbox/applied/65a686fe-7165-4fc9-9f1b-4b98c1824203.json similarity index 100% rename from docs/outstanding-issues-inbox/65a686fe-7165-4fc9-9f1b-4b98c1824203.json rename to docs/outstanding-issues-inbox/applied/65a686fe-7165-4fc9-9f1b-4b98c1824203.json diff --git a/docs/outstanding-issues-inbox/74e3183a-6e6a-49ef-9e8d-6bb329396ca6.json b/docs/outstanding-issues-inbox/applied/74e3183a-6e6a-49ef-9e8d-6bb329396ca6.json similarity index 100% rename from docs/outstanding-issues-inbox/74e3183a-6e6a-49ef-9e8d-6bb329396ca6.json rename to docs/outstanding-issues-inbox/applied/74e3183a-6e6a-49ef-9e8d-6bb329396ca6.json diff --git a/docs/outstanding-issues-inbox/7c385a4f-f54b-41f8-8841-883cd1995068.json b/docs/outstanding-issues-inbox/applied/7c385a4f-f54b-41f8-8841-883cd1995068.json similarity index 100% rename from docs/outstanding-issues-inbox/7c385a4f-f54b-41f8-8841-883cd1995068.json rename to docs/outstanding-issues-inbox/applied/7c385a4f-f54b-41f8-8841-883cd1995068.json diff --git a/docs/outstanding-issues-inbox/81b9bb52-9f07-496e-b12b-7795c93c8e24.json b/docs/outstanding-issues-inbox/applied/81b9bb52-9f07-496e-b12b-7795c93c8e24.json similarity index 100% rename from docs/outstanding-issues-inbox/81b9bb52-9f07-496e-b12b-7795c93c8e24.json rename to docs/outstanding-issues-inbox/applied/81b9bb52-9f07-496e-b12b-7795c93c8e24.json diff --git a/docs/outstanding-issues-inbox/8da919a5-4e1b-417c-89df-7a1692a61f17.json b/docs/outstanding-issues-inbox/applied/8da919a5-4e1b-417c-89df-7a1692a61f17.json similarity index 100% rename from docs/outstanding-issues-inbox/8da919a5-4e1b-417c-89df-7a1692a61f17.json rename to docs/outstanding-issues-inbox/applied/8da919a5-4e1b-417c-89df-7a1692a61f17.json diff --git a/docs/outstanding-issues-inbox/a56255bc-102a-4483-b7eb-c387e758054e.json b/docs/outstanding-issues-inbox/applied/a56255bc-102a-4483-b7eb-c387e758054e.json similarity index 100% rename from docs/outstanding-issues-inbox/a56255bc-102a-4483-b7eb-c387e758054e.json rename to docs/outstanding-issues-inbox/applied/a56255bc-102a-4483-b7eb-c387e758054e.json diff --git a/docs/outstanding-issues-inbox/a718d880-9295-4019-8a60-d0952293668b.json b/docs/outstanding-issues-inbox/applied/a718d880-9295-4019-8a60-d0952293668b.json similarity index 100% rename from docs/outstanding-issues-inbox/a718d880-9295-4019-8a60-d0952293668b.json rename to docs/outstanding-issues-inbox/applied/a718d880-9295-4019-8a60-d0952293668b.json diff --git a/docs/outstanding-issues-inbox/ae71cb9d-8094-44aa-bb1e-f1022156acc5.json b/docs/outstanding-issues-inbox/applied/ae71cb9d-8094-44aa-bb1e-f1022156acc5.json similarity index 100% rename from docs/outstanding-issues-inbox/ae71cb9d-8094-44aa-bb1e-f1022156acc5.json rename to docs/outstanding-issues-inbox/applied/ae71cb9d-8094-44aa-bb1e-f1022156acc5.json diff --git a/docs/outstanding-issues-inbox/b9a9492c-5cb5-4e13-8428-fbb3f70d979e.json b/docs/outstanding-issues-inbox/applied/b9a9492c-5cb5-4e13-8428-fbb3f70d979e.json similarity index 100% rename from docs/outstanding-issues-inbox/b9a9492c-5cb5-4e13-8428-fbb3f70d979e.json rename to docs/outstanding-issues-inbox/applied/b9a9492c-5cb5-4e13-8428-fbb3f70d979e.json diff --git a/docs/outstanding-issues-inbox/b9aea308-80d5-43e1-b464-902e499e46a4.json b/docs/outstanding-issues-inbox/applied/b9aea308-80d5-43e1-b464-902e499e46a4.json similarity index 100% rename from docs/outstanding-issues-inbox/b9aea308-80d5-43e1-b464-902e499e46a4.json rename to docs/outstanding-issues-inbox/applied/b9aea308-80d5-43e1-b464-902e499e46a4.json diff --git a/docs/outstanding-issues-inbox/bd007515-51b0-49d5-ad9e-4e97ea701ab4.json b/docs/outstanding-issues-inbox/applied/bd007515-51b0-49d5-ad9e-4e97ea701ab4.json similarity index 100% rename from docs/outstanding-issues-inbox/bd007515-51b0-49d5-ad9e-4e97ea701ab4.json rename to docs/outstanding-issues-inbox/applied/bd007515-51b0-49d5-ad9e-4e97ea701ab4.json diff --git a/docs/outstanding-issues-inbox/bd931cc7-71ad-430e-8921-241633e1c5be.json b/docs/outstanding-issues-inbox/applied/bd931cc7-71ad-430e-8921-241633e1c5be.json similarity index 100% rename from docs/outstanding-issues-inbox/bd931cc7-71ad-430e-8921-241633e1c5be.json rename to docs/outstanding-issues-inbox/applied/bd931cc7-71ad-430e-8921-241633e1c5be.json diff --git a/docs/outstanding-issues-inbox/ca09fd31-c3c7-43f1-8277-34afa27c85c6.json b/docs/outstanding-issues-inbox/applied/ca09fd31-c3c7-43f1-8277-34afa27c85c6.json similarity index 100% rename from docs/outstanding-issues-inbox/ca09fd31-c3c7-43f1-8277-34afa27c85c6.json rename to docs/outstanding-issues-inbox/applied/ca09fd31-c3c7-43f1-8277-34afa27c85c6.json diff --git a/docs/outstanding-issues-inbox/d0a9a345-d7da-44e4-b512-e7d413ad2239.json b/docs/outstanding-issues-inbox/applied/d0a9a345-d7da-44e4-b512-e7d413ad2239.json similarity index 100% rename from docs/outstanding-issues-inbox/d0a9a345-d7da-44e4-b512-e7d413ad2239.json rename to docs/outstanding-issues-inbox/applied/d0a9a345-d7da-44e4-b512-e7d413ad2239.json diff --git a/docs/outstanding-issues-inbox/da3fce96-1480-40e9-b2fd-07dcb7d3bae8.json b/docs/outstanding-issues-inbox/applied/da3fce96-1480-40e9-b2fd-07dcb7d3bae8.json similarity index 100% rename from docs/outstanding-issues-inbox/da3fce96-1480-40e9-b2fd-07dcb7d3bae8.json rename to docs/outstanding-issues-inbox/applied/da3fce96-1480-40e9-b2fd-07dcb7d3bae8.json diff --git a/docs/outstanding-issues-inbox/e9386425-20e9-436d-ab13-7e076310c5bb.json b/docs/outstanding-issues-inbox/applied/e9386425-20e9-436d-ab13-7e076310c5bb.json similarity index 100% rename from docs/outstanding-issues-inbox/e9386425-20e9-436d-ab13-7e076310c5bb.json rename to docs/outstanding-issues-inbox/applied/e9386425-20e9-436d-ab13-7e076310c5bb.json diff --git a/docs/outstanding-issues-inbox/f7499c28-d1ac-4230-abbe-517b02001f7e.json b/docs/outstanding-issues-inbox/applied/f7499c28-d1ac-4230-abbe-517b02001f7e.json similarity index 100% rename from docs/outstanding-issues-inbox/f7499c28-d1ac-4230-abbe-517b02001f7e.json rename to docs/outstanding-issues-inbox/applied/f7499c28-d1ac-4230-abbe-517b02001f7e.json diff --git a/docs/outstanding-issues-inbox/fc70be6e-e56f-4377-99b6-54745c5c5ea6.json b/docs/outstanding-issues-inbox/applied/fc70be6e-e56f-4377-99b6-54745c5c5ea6.json similarity index 100% rename from docs/outstanding-issues-inbox/fc70be6e-e56f-4377-99b6-54745c5c5ea6.json rename to docs/outstanding-issues-inbox/applied/fc70be6e-e56f-4377-99b6-54745c5c5ea6.json diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index 865914dc9c..6defda7b1d 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -91,46 +91,36 @@ removed after current-main verification; it is not missing recommended work. | #018 | P2 | task | Split the lithium, ADHD and metabolic residuals by mechanism | Current evidence keeps the mechanisms separate. **Lithium — closed within this item:** the row/atom-aware subject guard, foreign-parameter rejection and query-specific range promotion returned `0.5–1.0 mmol/L` with correct targeting/citation; the full retrieval canary remained 36/36 with recall 1.0 and zero per-case RR regressions, and the full answer canary passed every blocking gate. **ADHD — open corpus debt:** `CG.MHSP.ADHD.pdf` is absent from the hosted corpus and the retrieved chart exposes `accessible_table_count=0`; repair corpus/fixture or ingestion evidence rather than weakening extractive budgets. **Metabolic — open structured-evidence debt:** the standalone plural classifier worsened the live answer and was reverted; obtain auditable schedule text/table evidence before another candidate. | targeted live lithium/ADHD/metabolic evidence 2026-07-27; `docs/evidence/rag-reliability-evidence-2026-07-27.md`; refuted approaches | 2026-07-21 | | #023 | P2 | task | Complete scheduled browser and labeling disposition | **Partial 2026-07-30:** `release-browser-matrix` no longer depends on `pr-required`, so a blocking scheduled dependency audit cannot skip Firefox/WebKit. Still need one green matrix datapoint + human irrelevant-at-10 disposition. The 2026-07-26 retrieval and answer artifacts are read and compared under resolved #051. Scheduled CI run `30216361999` failed its existing production dependency audit before Firefox/WebKit, while production Chromium passed. After that audit is green, capture one scheduled/manual browser-matrix datapoint; separately record the human decision for the stable irrelevant-at-10 set. #084 now makes each top-10 grade and matched signal reproducible, but it does not substitute for the human disposition. Do not rerun or spend on RAG for this item. | runs `30216191889`/`30216361999`; per-rank diagnostics #084; session 2026-07-27 | 2026-07-21 | | #099 | P2 | task | Remove the remaining fixed per-request round trips | **Outcome:** the answer path stops paying avoidable per-request Supabase round trips. **Done 2026-07-29:** shared-cache-hit promotion deferred off the response path with its mid-request staleness guard intact and documented (`rag.ts:3234`, `rag-cache.ts`); scope resolution overlapped with the rate-limit RPC, signal threaded so a client disconnect finally cancels its paginated queries (`answer/route.ts`). **REFUTED on PR #1377 review — do not retry:** the same pass also overlapped scope with the rate-limit RPC and aborted it on deny, claiming the limiter could "deny for free". It cannot. With caller-supplied `filters` or explicit ids, scope passes its zero-query early returns (`search-scope.ts:242,253`) into the paginated `documents` loop at `:269`, and an `AbortSignal` cancels the client request without un-executing a statement Postgres already began — so throttled traffic kept burning database capacity while collecting 429s, against `capacity-review.md:106-113`'s first-soft-failure warning. Scope is behind admission again, pinned by `tests/answer-route-preamble.test.ts`. Re-attempting the overlap requires a non-database admission gate ahead of the durable limiter first. **Remaining:** (a) the 8 `setCachedSearch` awaits — deferring changes `throwIfAborted` semantics and widens a real mutation window because the clone happens after an `await`, so each branch needs discharging individually; (b) batch the anonymous subject+global rate-limit pair, which needs a NEW atomic RPC modelled on `consume_summary_rate_limits_atomic` and cannot be called until the operator applies it — `Promise.all` is the WRONG fix because it consumes the global bucket even when the subject bucket already denied; (c) stop the proxy and route handler resolving identity twice per authenticated request — no in-process memo can do this (different `Request` objects), so the proxy must forward unspoofable verified claims via a header it controls. Cross-references #011: halving auth resolutions eases the ~10-connection Auth cap that `capacity-review.md:106-113` calls the first hard failure. | `docs/audit/latency-audit-2026-07-28.md` L1-1/L1-3/L1-4; `src/lib/api-rate-limit.ts:276-282`; `src/proxy.ts:125` | 2026-07-29 | -| #100 | P2 | rec | Buffered answer generation has no incremental verified delivery | UPDATE 2026-08-13 (PR #1909): Phase 0 offline contract proof and flag-gated Phase 1 server emission implemented (RAG_INCREMENTAL_EVIDENCE_PREVIEW, default false). Remaining: client parsing/rendering phase behind its own flag + verify:ui, then the design's provider-backed acceptance gates before production enablement; Phase 2 stays provider-gated. **Design complete; runtime work remains provider-gated.** [`verified-answer-incremental-delivery-design.md`](verified-answer-incremental-delivery-design.md) records the clinical-governance decision and staged contract: keep the `progress`/`final`/`error` allowlist; disclose bounded, owner-scoped evidence only after the canonical danger-level source-governance refusal permits it, then emit complete answer sections only after each reuses the full production verification boundary; reconcile every preview byte-for-byte with the authoritative `final`; discard all previews on error/cancel/retry; deploy behind separate parse/emission/render flags. Phase 0 contract proof and Phase 1 evidence preview can be developed offline, but visible rollout still needs clinical/browser proof. Phase 2 changes generation architecture and requires explicit approval for answer-quality evals plus a baseline/post live canary pair. **Naive token streaming remains REFUTED:** never re-land `token`, `revising`, provisional prose, or a weaker stream-only verifier. Cross-references #021. | `docs/verified-answer-incremental-delivery-design.md`; `docs/audit/latency-audit-2026-07-28.md` L0-1; `src/lib/answer-stream-contract.ts:18-21` | 2026-07-30 | -| #102 | P3 | task | Apply the additive `documents` index debt (operator) | **Outcome:** bare-column `ILIKE` and the paged status scan on `documents` are index-served on hosted. `documents_title_trgm_idx` indexes a CONCATENATED expression, so the bare-column predicates in `api/documents/route.ts:193` and `rag-candidate-sources.ts:477` (RAG path) cannot use it and fall back to scanning; `search-scope.ts:271-277` sorts per page against the single-column `documents_status_idx`. **Runbook prepared 2026-07-29 — NOT applied, item stays open:** three `CREATE INDEX CONCURRENTLY` statements authored and reviewed in `docs/operator-apply-performance-latency-remediation.md` — additive, though **the "recall is byte-identical" claim was RETRACTED on 2026-07-29 review**: `fetchDocumentTitleAliasRows` (`rag-candidate-sources.ts:482`) applies `.limit(12)` with no `ORDER BY`, so a new index can change which title-alias documents feed candidate assembly. Only the documents-list use stays ordering-safe; `(status,id)` is canary-gated too — see runbook, and making that `.limit(12)` deterministic first does **not** lift the gate — an unordered `LIMIT` has no stable selection to preserve, so imposing an order can pick a different twelve and is itself an ordering behaviour change on a retrieval surface, which AGENTS.md requires a canary pair for. Sequencing the ordering fix first is worthwhile (unordered `LIMIT` on a retrieval input is latent nondeterminism regardless) but yields two canary-gated changes, not one (PR #1377 review). **Deliberately NO migration file:** an additive-index migration without a synchronized `schema.sql` mirror and regenerated drift manifest is exactly what closed PR #1312, and the mirror cannot come first because `required_indexes` in `search_schema_health()` (`schema.sql:3178`) runs against live. **Next (operator):** **author the migration first** — `supabase/migrations/` is the source of truth and `schema.sql` only a mirror, so hand-run operator SQL never reaches staging, disaster-recovery replay, or a local `supabase db reset`, and a `required_indexes` registration would fail there (PR #1377 review); follow the `20260717170000_registry_projection_cleanup.sql` idempotent pattern. **That migration must also carry the health-function change** — `required_indexes` lives inside `search_schema_health()`, which is redefined by `create or replace function` in eleven migrations (copy `20260705180000_reconcile_search_health_indexes.sql:62`); editing `schema.sql:3177` alone moves only the mirror and leaves the indexes unmonitored on hosted (PR #1377 review). Then apply concurrently, confirm `indisvalid`, mirror both the index statements and the identical function body into `schema.sql`, run `npm run drift:manifest` (Docker), and deploy the migration LAST — in that order, in one change. Expect `check:drift` to report them as unexpected between steps 1 and 2. **Rollback is three deployed phases, not the reverse of one:** retract `required_indexes` via its own `create or replace function` migration and deploy → drop concurrently live → only then deploy the `schema.sql` removal plus an idempotent forward `drop index if exists` migration, because Supabase wraps migrations in a transaction and a plain `DROP INDEX` there takes the lock the concurrent procedure exists to avoid (PR #1377 review). | `docs/audit/latency-audit-2026-07-28.md` L2-3/L2-5; `docs/operator-apply-performance-latency-remediation.md` | 2026-07-29 | +| #100 | P2 | rec | Buffered answer generation has no incremental verified delivery | UPDATE 2026-08-21 (repo read on main at 1cc0d2987, no provider access): the client-side half has landed. NEXT_PUBLIC_RAG_INCREMENTAL_EVIDENCE_PREVIEW_RENDER is present in .env.example (commented, default false) and is consumed by src/lib/client-env.ts, so the Phase 1 client parsing/rendering flag exists alongside the server-side RAG_INCREMENTAL_EVIDENCE_PREVIEW=false. Remaining scope is therefore narrower than recorded: verify:ui proof of the client render path, then the design's provider-backed acceptance gates before production enablement. Phase 2 stays provider-gated. Not verified: whether the render path is actually exercised by a UI journey. | `docs/verified-answer-incremental-delivery-design.md`; `docs/audit/latency-audit-2026-07-28.md` L0-1; `src/lib/answer-stream-contract.ts:18-21` | 2026-07-30 | +| #102 | P3 | task | Apply the additive `documents` index debt (operator) | UPDATE 2026-08-21 (read-only Supabase MCP get_advisors performance lint against production ref sjrfecxgysukkwxsowpy): documents_title_trgm_idx exists on public.documents and is reported by the unused_index lint as never used. TREAT THAT AS WEAK EVIDENCE, NOT CONFIRMATION: 20260819100200_restore_search_health_trigram_indexes was applied two days earlier and recreating an index resets its usage statistics, so a zero-use reading is expected regardless of whether the bare-column ILIKE predicates can reach it. The same lint currently reports 31 unused indexes, several of them freshly restored in the 20260819100000-100300 batch, which is consistent with a stats reset rather than dead indexing. The row's actual claim - that the index covers a CONCATENATED expression and so cannot serve the bare-column predicates in the documents API route and rag-candidate-sources - was NOT tested, because that needs EXPLAIN or a pg_indexes read and SQL execution was blocked in this session. Re-measure with EXPLAIN in the operator window before applying the prepared runbook. | `docs/audit/latency-audit-2026-07-28.md` L2-3/L2-5; `docs/operator-apply-performance-latency-remediation.md` | 2026-07-29 | | #191 | P3 | task | X5: ACL-migration consolidation (provider-gated) | **Outcome:** ACL-related migrations are consolidated per maturity work-order X5 without weakening owner-scope/RLS. **Next:** DB-owner approved window only; live-DB provider confirmation required before apply. **Stop:** no hosted apply from an agent session without explicit approval. | docs/maturity-backlog-workorders.md X5; #086 | 2026-07-31 | | #231 | P2 | issue | Generation fallbacks no longer stick in answer cache; lithium generation quality still falls back safely | S1 (#2022), S1b (#2035), S1d (#2054) landed with green canary pairs: generation-quality false rejections and the finalizer gap hole are fixed. Residual R4: chronic ~30 s strong-route provider_timeout on metformin-renal-dosing and valproate-pregnancy (measured 4/4 at v18 4ea310e48 and 2/4 at v19 main on 2026-08-18; safe source-backed extractive fallback, never model synthesis). Route-budget stop condition unchanged. Remediation-plan Phase 5.2 (re-test #231 after the trigram-index restore) is satisfied by S1's 2026-08-17 healthy-latency probes; live-drift #316 Phase 1.2 classified all ten RPC divergences as attribute-only, so canaries since 2026-08-14 measure a reconstructable path. Re-graded P1 -> P2. Next: prompt/context trimming for complex classes, or accept the extractive fallback as the durable answer for these two classes. Do not add a separate R4 row — this update supersedes that need. | docs/rag-improvement/HANDOVER.md secs 1-2; docs/rag-improvement/COORDINATION.md sec 7; docs/database-remediation-plan.md Phase 5.2; docs/audit/live-drift-forensics-2026-08.md sec on Phase 1.2 classification; ledger rows #316 / #248 / #231 / #342; session 2026-08-18 | 2026-08-04 | | #308 | P3 | issue | Desktop /documents/search CLS is 0.119, above threshold and stable across runs and baselines | Measured 2026-08-12 during the #147 close-out, twice, on the offline Lighthouse harness (Chromium 141): desktop /documents/search CLS **0.119**, against a committed baseline that also reads **0.119**. So this is long-standing and deterministic, not a regression — and it is above the 0.1 threshold. It sits outside #147's scope, which was mobile only, and it contradicts that row's claim that 'desktop passes everywhere: 0.016-0.097' — that range is stale. Companion desktop values from the same runs, all passing: /dsm 0.014, /forms 0.059-0.064, / 0.006, /therapy-compass 0.000. Desktop attribution completed 2026-08-14: a Playwright + PerformanceObserver(layout-shift) harness against an offline production build at 1350x940 DPR 1 recorded **0.118** CLS. This is a separate attribution measurement, not a replacement for the canonical 0.119 Lighthouse value. One first-paint+~0.3-0.5s event contributed ~99.98% of that harness total: MasterSearchHeader's composer-adoption effect portals the search composer into GlobalSearchShell's desktop slot, while the header shrinks 184px and the slot grows 0 -> 184px. This is shared desktop search-chrome timing, not page-local. Next: reserve the settled height at the adoption boundary under the one-composer/hidden-means-zero-reserve contracts, then re-measure with the same harness. Stop: do not raise the CLS budget; do not read local LCP or TBT from the loopback harness; and do not use a blanket min-height that hides the shift without matching the header reserve. | Local offline verify:lighthouse runs 2026-08-12 (two runs, identical CLS); #147 close-out; lighthouse-budget.json. Attribution: session 2026-08-14, PR branch codex/visual-layout-polish; desktop CLS script adapted from scripts/measure-cls-attribution.mjs (offline, not committed). | 2026-08-12 | | #316 | P1 | issue | Live DB has 20 currently missing repo-defined indexes and 10 retrieval RPC bodies diverge; weekly live-drift has been red since 2026-07-26 with no routing | PHASE 4 COMPLETE 2026-08-19; the RPC track of this row closed 2026-08-18. BOTH HALVES OF THIS ROW ARE NOW CLOSED. (A) RPC divergence, from the 2026-08-18 window (PR #2123, forensics 3.7): the authorised window's pre-flight found the pending set EMPTY -- all five 20260818 migrations were already applied with executed statements (3/11/12/5/4, the CLI db push shape, not mark-applied) -- so db push was never run, no migration repair, no vault reads, zero production writes. Manifest def_hash equals live for all ten match_* functions, and live-drift 32131517648 showed 0 function mismatches. CAUSE, platform finding since resolved: the Supabase GitHub integration (Branching, production bound to git main, branch record 2026-06-27) was auto-applying every migration merged to main -- push-triggered live-drift bracketed 110000-112000 to 34 s after #2106's squash-merge. D4 IS NOW DECIDED: auto-deploy is OFF, confirmed empirically on 2026-08-19 when the four new 20260819 migrations sat pending on production after the branch existed and reached it only via an explicit db push. Not established: who enabled the integration or when, or whether the July mark-applied rows trace to it. (B) Index restoration, 2026-08-19 owner-authorized off-peak window: all 20 missing_live indexes rebuilt with CREATE INDEX CONCURRENTLY from canonical definitions cross-read against their defining migrations -- Batch A 14/14, Batch B 6/6, every one indisvalid AND indisready with normalized pg_get_indexdef matching canonical, zero invalid builds, zero retries, zero skips, zero lock waits (pg_locks read between every Batch B build). No transactional build was ever attempted; #102's bare-column indexes held out. Both unexpected_live indexes were DROPPED CONCURRENTLY rather than codified, because the chain already commands both drops and each is a strict leading-column subset of a present canonical index: document_table_facts_document_id_idx (superseded per 20260620000000) and storage_cleanup_jobs_owner_id_idx (superseded per 20260703030000/20260708000000). Live now reports 210 public indexes against the manifest's 210, zero invalid anywhere. Codified in five migrations applied by real supabase db push (never migration repair; every history row carries executed statements): 20260819100000/100100 guard Batch A/B, 20260819100200 discharges the plan 4.4 debt by guarding the two trigram indexes restored 2026-08-14 that 20260804110240 never checked, 20260819100300 takes search_schema_health() required_indexes 22->30 adopting all 8 Phase 6.3 monitor-candidates (unmonitored list 44->36, no monitor-candidate left; production ok true), and 20260819100150 repairs a chain defect the guard itself caught -- see below. Live-drift 32171070287: UNEXPECTED DRIFT 37->16, missing_live 20->ZERO, unexpected_live 2->ZERO. Staging brought to full parity in the same task; its drift comparison is GREEN with ZERO unexpected drift (was 19), corpus untouched, --prune-stale correctly not used. THE GUARD EARNED ITS KEEP: 20260819100200 failed the Supabase Preview check on PR #2151 because a preview branch builds from the migration chain alone, and the chain permanently produced the WRONG document_chunks_content_trgm_idx -- 20260606000000 creates it first without coalesce(content,''), and both later correct creators use IF NOT EXISTS so they no-op, with no migration ever dropping it. Forensics 3.3(d) had scoped this as staging-only and hand-repaired it there; it was never staging-only (db reset, DR replay, CI migration replay, preview branches all get the wrong index, which is NULL for rows with NULL content and so silently omits those chunks). Fixed by 20260819100150, conditional so it no-ops when canonical, rebuilds only on an empty table, and raises rather than run a write-blocking build on a populated one; proven by replaying the whole chain into a scratch Postgres (fails without it exactly as CI did, 204/204 with it) and the no-op path proven on production itself (index OID unchanged at 1491258 across the push). TWO ESCALATIONS FOR THE OWNER, neither absorbed. (1) PITR IS NOT ENABLED on production (pitr_enabled false, walg_enabled true, daily physical backups only, latest 2026-08-17T20:33:28Z), so the plan's standing 'restore point before any mutating phase' rule cannot be met; Phase 4 proceeded only because every statement was index-only with an exact one-statement inverse, and no future window that mutates DATA should proceed on that precedent. Queued separately as its own P2. (2) The migration_history block did NOT drop and no allowlist entry was written -- measured, not skipped: of the 15 no-statements versions, 6 are index-shaped and the intersection between the objects they create and the 22 these guards validate is EMPTY (near-misses are distinct objects, e.g. audit_logs_owner_id_idx vs audit_logs_owner_created_idx). The 15 stay unallowlisted and remain #Q5JHBJ's work. REMAINING FOR THIS ROW: nothing on the index or RPC tracks. Phase 5 measurement (after-EXPLAIN set, #231 re-test on healthy latency, check:production-readiness) is the only follow-on. Full evidence with dates, run IDs and pasted output in docs/audit/live-drift-forensics-2026-08.md sections 3.7 and 'Phase 4 completion'. Session traps still current: the main checkout D:\Repos\Database is linked to STAGING, so link a dedicated worktree for production and unlink after; supabase db query --linked --project-ref works read-only via the management API without a DB password; db query parses a leading -- as a flag, so pass SQL that starts with a comment via --file; production has no track_commit_timestamp. | PR #2151 (Phase 4) and PR #2123 (window 3.7); live-drift runs 32171070287 and 32131517648; forensics sections 3.7 and 'Phase 4 completion' | 2026-08-13 | | #321 | P3 | task | Four follow-up groups cover nine controls after #291 | PARTIAL 18 August 2026. Of the four follow-up groups: (1) the filmstrip 'Page unknown' control is FIXED — document-image-filmstrip.tsx converted its data-driven disabled state from native disabled to aria-disabled=true + ignoreUnavailableActivation + an sr-only reason, per docs/wiring-conventions.md's stated-reason pattern (settles this one control from #291's follow-up list); tests/document-image-filmstrip.dom.test.tsx gained a focused case (aria-disabled, not natively disabled, accessible description, click is a no-op), vitest run: 3 passed. The other three groups are unchanged and still not single-PR-sized: the six differential comparison page controls remain coupled to its own planned rewrite and pinned density test; DocumentViewer's persistent-access-reason/transient-loading split is a classification design decision, not yet made; the pin-limit control remains a capacity-state judgement call. Stays open for those three. | PR #1778 body; verified against main 2d27039 | 2026-08-14 | | #339 | P2 | task | Favourites Continue and Recent are driven by hard-coded demo timestamps; real saved items have no last-opened data | Surfaced while shipping #164 (PR #1983), which made both surfaces prominent. src/components/clinical-dashboard/favourites-command-library-page.tsx derives 'most recently used' from lastUsedScore(item.lastUsed), and item.lastUsed comes from lastUsedByItemId — a hard-coded five-entry literal keyed to demo slugs ('Today 08:44', 'Yesterday 16:12', ...). Anything else, including every real registry favourite, falls back to the literal string 'Saved', which lastUsedScore buckets at 1000. pinnedItemIds is likewise a hard-coded two-item Set. The consequence after #164: for a signed-in user with real favourites, the Continue card and the Recent panel are effectively arbitrary — every item ties at the same score and the order is whatever the source array happened to be. Note that recentQueries in the shell is search-query history, not viewed-item history, so it cannot back this. Next: add a per-favourite last-opened timestamp. Cheapest is a client-side recents store keyed by favourite id written on open; the durable version is a column on the account favourites record so it survives a device change, which is a schema plus /api/account/favourites change and needs the usual migration review. Either way, pinning should stop being a hard-coded id set. Stop: do not fabricate a timestamp at render time from anything other than a recorded open event — an invented 'last used' on a clinical reference list is worse than an honest absence. | session 2026-08-15; PR #1983; favourites-command-library-page.tsx lastUsedByItemId/pinnedItemIds | 2026-08-15 | -| #4TBHS8 | P3 | issue | Advisory UI mockup spec 'phone filter sheet follows the shared local-filter behavior' fails on main | tests/ui-tools-search-mode-mockup.spec.ts:188 fails at line 198 waiting for '2 showing' inside [data-testid=tools-search-filter-sheet] after searching 'Safety' at 390px. Reproduced locally under --project=chromium-mockups on BOTH claude/card-review-optimize-h0pidc and origin/main (dc7e518), so it is pre-existing and NOT caused by the card branch — attribution was checked before any fix was attempted. It surfaced now only because the ui-advisory lane fires on advisory_ui_changed (a mockup surface changed or the flake ledger is non-empty) and had been skipped on every earlier run of that PR. It is non-blocking: ui-advisory carries continue-on-error true and is absent from pr-required's needs list in ci.yml, and verify:ui excludes @mockup via --grep-invert, which is why a 429-pass local run never touched it. Next: open the trace at test-results/ui-tools-search-mode-mocku-b6163-hared-local-filter-behavior-chromium-mockups/trace.zip and decide whether the expected count of 2 is stale against the current tools catalogue or the facet hint genuinely miscounts; the mockup renders the production ToolsSearchResultsPage, so a real miscount would affect /tools too. Stop: do not change the expected number to match observed output without establishing which is correct. | session 2026-08-18; PR #2060 Advisory UI run 32090358678; reproduced on origin/main dc7e518 | 2026-08-18 | +| #4TBHS8 | P3 | issue | Advisory UI mockup spec 'phone filter sheet follows the shared local-filter behavior' fails on main | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): the specific failure mode recorded here - waiting for the exact text 2 showing inside the tools-search-filter-sheet testid - can no longer occur, because that string has zero occurrences in tests/ui-tools-search-mode-mockup.spec.ts. The test itself still exists and was rewritten. NOT VERIFIED: whether it passes. Re-run before closing. Tracked with #SZGPAH, which names the same file. | session 2026-08-18; PR #2060 Advisory UI run 32090358678; reproduced on origin/main dc7e518 | 2026-08-18 | | #2AB2NJ | P3 | task | Owner decision: enable RAG_TELEMETRY_EXTENDED (verification_latency_ms projection) in production once a dashboard consumer exists | Packet S5 (PR #2056, merge 093f9340c) landed the B1 telemetry gap assessment: the one proven gap is verification_latency_ms, now persisted behind RAG_TELEMETRY_EXTENDED (typed, default false) via the allow-listed projection module with canary-absence tests. Enabling it in production is an owner decision gated on a dashboard consumer existing (no consumer today), and is a Railway env change (provider-backed, explicit approval; rollback = set false). Next: when a dashboard question needs verification latency, set RAG_TELEMETRY_EXTENDED=true on the Database service after confirming the canary-absence tests are still green on main. Stop: do not enable speculatively; do not add unproven fields. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | | #C2D9JF | P2 | issue | Adversarial divergence (S5 harness pin): scope-other-owner-document — abstains in substance but the review fallback still cites in-scope evidence | Pinned in tests/rag-adversarial-harness.test.ts KNOWN_DIVERGENCES (self-expiring). Observed shape: grounded false, confidence unsupported, but cited chunk ids [syn-scope-owner-a] — the answer correctly abstains from the other-owner document, yet the review fallback attaches an in-scope citation to an unsupported answer. Fixture: scripts/fixtures/rag-adversarial-cases.v1.json case scope-other-owner-document (category scope_or_tenant). Tenancy/no-read invariant held (the other-owner content is never read). Next: decide whether an unsupported abstention may carry any citation; if not, strip citations on the abstention path (RAG-surface change; own PR; harness pin flips; canary pair). Stop: do not delete the pin without the behaviour change. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | -| #71NT23 | P2 | task | No mobile-WebKit or display-mode:standalone Playwright project exists; phone coverage is a narrow viewport on desktop engines | playwright.config.ts defines chromium, chromium-mockups, firefox and webkit, and the webkit project uses devices["Desktop Safari"]. Every phone assertion in the suite is therefore a narrow viewport on a desktop engine, and nothing exercises display-mode: standalone at all — despite globals.css:3570-3606 carrying the only phone-AND-standalone-exclusive CSS in the repo (bounded overflow:hidden shell, -webkit-overflow-scrolling:touch scrollport). PR #2046 added per-test devices["iPhone 14"] emulation in tests/ui-phone-motion.spec.ts as a cheap partial, but that is still Playwright WebKit, not the iOS engine. Next: decide between a dedicated mobile-WebKit project (CI cost) and per-test emulation as the standing pattern, and add standalone display-mode coverage for the phone shell rules. Relates to #280 (physical iPhone acceptance debt). | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | +| #71NT23 | P2 | task | No mobile-WebKit or display-mode:standalone Playwright project exists; phone coverage is a narrow viewport on desktop engines | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): partially advanced. playwright.config.ts now carries a documented phone-PWA standalone emulation strategy annotated with this issue id, and defines chromium, chromium-mockups, firefox and webkit projects. What still appears absent is a genuinely mobile-WebKit project - the webkit entry is the desktop engine - and no display-mode standalone project was found. Re-measure the exact project list before scoping the work; physical iPhone Safari/PWA acceptance remains out of scope for any container evidence. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | | #NTAV3D | P2 | issue | Adversarial divergence (S5 harness pin): scope-guessed-chunk-id — review fallback returns a grounded source pointer echoing the query instead of refusing | Pinned in tests/rag-adversarial-harness.test.ts KNOWN_DIVERGENCES (self-expiring). Observed shape: grounded true, cited [syn-scope-guess-a], and the guessed (never-retrieved) chunk id syn-not-retrieved-zzz is never resolved into content — the no-read invariant holds — but the review fallback returns a grounded source pointer that echoes the query text rather than refusing the guessed-id request. Fixture: scripts/fixtures/rag-adversarial-cases.v1.json case scope-guessed-chunk-id (category scope_or_tenant). Next: decide whether a query naming an unretrieved chunk id should refuse rather than fall back to a source pointer (RAG-surface change; own PR; harness pin flips; canary pair). Stop: do not delete the pin without the behaviour change. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | | #VXB8XA | P2 | issue | Adversarial divergence (S5 harness pin): cite-mismatched-attribution — offline document-match listing cites every retrieved document, not only the claim-bearing one | Pinned in tests/rag-adversarial-harness.test.ts KNOWN_DIVERGENCES (self-expiring: the harness asserts the normative B0 fixture contract still FAILS; when behaviour reaches the fixture expectation the pin goes red and must be deleted). Observed shape: cited chunk ids [syn-cite-attrib-a, syn-cite-attrib-b], grounded true, answerQualityTier source_only — the document-match listing attributes both retrieved documents although only one carries the claim. Fixture: scripts/fixtures/rag-adversarial-cases.v1.json case cite-mismatched-attribution (category citation_fabrication). Safety invariants (network, budget, canary absence, forbidden substrings, tenancy) hold; only citation attribution precision diverges. Next: decide whether the document-match listing should cite only claim-bearing documents (RAG-surface change; own PR; RAG impact line; offline harness proves the pin flips; canary pair). Stop: do not delete the pin without the behaviour change; do not weaken the fixture. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | | #S4K1GA | P3 | task | Physical iPhone acceptance owed for the answer-progress motion fix (Safari + installed PWA, Motion=Full) | PR #2046 fixed the reported defect (OS Reduce Motion froze every animation and set the ECG trace to opacity:0) and added a Motion preference whose "full" value opts back in over the OS setting. All executed browser evidence ran on Chromium 1194 in a Cloud container — the repo's own verify:ui gate could not run because check:playwright-browser-revision reports the known #255 drift (expects 1234). Playwright WebKit is not the iOS engine either. Acceptance: on the physical iPhone, in Safari and as the installed PWA, with Settings > Motion set to Full, confirm the ECG strip visibly travels and the current-step spinner rotates; with Motion left on System, confirm the trace stays visible and static rather than blank. Failure to confirm means the defect class is unclosed, which is exactly how #1974/#1989/#1995 were each declared fixed. Relates to #255, #280. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | -| #Q5JHBJ | P2 | task | Deploy the 20260818090000 schema_drift_snapshot v2 history probe (Phase 6.1) in an approved production window after Phase 4, triage the first unguarded no-statements report, and author fail-fast guard migrations for the pre-contract 2026-07-01..02 and 2026-07-12 rows | PHASE 6.2 COMPLETE 2026-08-19 (owner-authorised production window). All fifteen no-statements versions classified and guarded with class validation (none earned superseded or no_ddl: no single later executed migration re-creates every object, and COMMENT ON is a catalog write, not an empty file). Six fail-fast guard migrations per 20260804110240: 20260819110000 dropped objects (absence of 7 functions + 4 indexes), 110100 catalog comments + purge-rag-retrieval-logs cron, 110200 three document FKs, 110300 forty-six operational index shapes, 110400 the index_generation_id promotion (6 columns + 6 indexes + 3 def_hashes), 110500 fifteen function def_hashes read from schema_drift_snapshot() itself. Allowlist now 20 entries (5 superseded + 15 validation). PROOF: full chain replay into the scratch image Applied 210/210 and CHAIN == MANIFEST (zero unexpected drift) -- every hand repair in sections 2.3/3.3/Phase 4 is reproduced by the chain, no reconcile migration needed; seven mutants raise and name their object; all six dry-ran green on production and a mutant fails there naming reset_document_index def_hash; production push real (migration list pending = exactly six, db push applied, rows stmt_count 4 no_statements false, 210 rows, documents 2851 untouched); staging by the Phase 2 method, six md5-matched rows, 210 rows, no_statements 0, drift comparison green. LIVE-DRIFT run 32251326536 on the branch: Compare step SUCCESS, all 20 history rows allowed, No unexpected schema drift -- #316's finding set is EMPTY for the first time since 2026-07-26. RESIDUAL, not this row's: the job still concludes failure because the Phase 0 step Align migration history (check:migration-history) ran for the first time ever and cannot read supabase_migrations over PostgREST (PGRST106, schema not exposed) -- queued as its own item. Evidence: forensics section '6.2 completion'. | PR #2123 forensics section 3.7; coordination chat 2026-08-19 | 2026-08-17 | -| #43SSS0 | P3 | rec | Three spring easing tokens in globals.css are dead: zero var() references and zero utility usage | --spring-tight, --spring-bouncy and --spring-gentle (src/app/globals.css:222-224) are declared in the @theme block but have no var() consumer in any stylesheet and no generated-utility consumer in src/. Tailwind v4.3.3 tree-shakes unused theme variables, so they never reach the compiled CSS — they are source noise, not shipped weight. Found while confirming (during PR #2046) that --animate-answer-ecg survives that same tree-shaking because it IS referenced via var() from the project's own CSS; --ease-spring is the working precedent for that pattern. Next: delete the three tokens, or wire them to the motion surfaces they were intended for. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | | #75JA0P | P2 | issue | Playwright runs the whole suite with reducedMotion:"reduce", so no gate reflects the default user configuration | playwright.config.ts:61 sets contextOptions: { reducedMotion: "reduce" } suite-wide, and every motion assertion has to opt out per-test via page.emulateMedia({ reducedMotion: "no-preference" }). That inversion is why three consecutive PRs (#1974, #1989, #1995) shipped green while a physical iPhone with OS Reduce Motion on showed a frozen, blank answer-progress panel: the suite never exercised the reported configuration. PR #2046 added tests/ui-phone-motion.spec.ts to cover that one surface, but the suite-wide default remains inverted for every other motion behaviour. Next: decide whether the suite default should be no-preference with reduce opted into per-test (the safer direction), or keep the current default and add a contract test that fails when a motion assertion has no explicit emulateMedia call. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | | #1PN5BM | P3 | issue | H5a residual: whether a constant similarity of 1 may contribute to a confidence label is still open, and after G1 it lives only in the hazard doc | Packet G1 (PR #2053, merged 2026-08-17) implemented owner decision Option B: buildDocumentSummaryResults now stamps similarity_origin "document_context" on document-summary rows, deriveConfidence is unchanged, and document summaries still reach "high". That closed the LEGIBILITY half of the H5a live residual -- the fabricated 1.0 is no longer indistinguishable from a perfect cosine at any surface that reads a row. It did NOT answer the underlying governance question: may a score nobody measured contribute to the confidence label a clinician reads at all? Option B was chosen because tagging has no measured safety cost while Option A (tag as synthetic_text, capping summaries at "medium") is a label downgrade without measured gain -- so the question was deferred deliberately, not resolved. The paired question row #J912J9 is being closed by G1, so once that closure reconciles this knowledge survives only in docs/clinical-hazard-analysis.md H5a and not in the queue anyone reads. NEXT: no action required unless a measured signal appears; if it does, the tag is what makes the fix cheap -- any future gate can now discriminate the document-summary route without re-deriving provenance. Guard rails already in place: tests/rag-score.test.ts pins the discriminating pair (two document_context citations >= 0.82 -> "high"; the identical scores tagged synthetic_text -> "medium"), so a silent change in either direction goes red. | Packet G1 session 2026-08-17 (PR #2053); docs/clinical-hazard-analysis.md H5a; closes-with #J912J9 | 2026-08-18 | -| #DREDWA | P3 | rec | Ledger writer self-tests use only legacy numeric ids, which is why a Crockford-id lookup bug survived the ULID migration unnoticed | Fixed in PR #2053: issueRowFingerprint (scripts/check-outstanding-issues.mjs) matched only /^#(\d+)$/ and keyed on entry.number, which is null on ULID-backed rows, so it returned null for every Crockford display locator -- and ledger-inbox.mjs reads a null fingerprint as "no such row" and refuses. npm run issues:done and issues:update were therefore unusable for EVERY row minted since the ULID migration, failing with "#J912J9 is not in Open items" about a row plainly in Open items. It surfaced only because a session happened to need to close two Crockford-id rows. ROOT CAUSE OF THE SURVIVAL, not of the bug: the self-tests and fixtures in scripts/outstanding-issues.mjs and tests/outstanding-issues-writer.test.ts exercise the writer almost entirely with legacy #005/#006/#013-style ids, so no test ever drove a Crockford id through the fingerprint path. A second, subtler trap sits in the same area and is now pinned but not generally guarded: Crockford's alphabet includes 0-9, so a ULID-derived locator can be ENTIRELY digits (the writer test's own id is #041061) and is indistinguishable from a legacy id by pattern -- branching on id shape rather than resolving against the table silently misses exactly those rows, which is how the first attempt at the fix still returned null. NEXT: add a fixture row with a ULID/Crockford id (ideally an all-digit one) to the shared ledger test fixtures and drive every writer entry point -- addIssue, resolveIssue, updateIssue, issueRowFingerprint, and the ledger-inbox done/update/reconcile paths -- through both id generations, so the next lookup left behind by an id-scheme change fails a test instead of a user's command. | Packet G1 session 2026-08-17 (PR #2053), discovered while queueing the G1 closures | 2026-08-18 | -| #90EVWZ | P3 | rec | check:drift clips table column diffs to 240 chars per side, so a wide-table column drift never names the column | scripts/check-drift.ts:192 (fieldDiff) serialises the whole alphabetised columns array of a table and clips each side to 240 characters, so for a wide table such as document_chunks (19 columns, several kB) a single late-alphabet column drift (token_estimate, forensics Phase 3 §3.3) prints two identical prefixes and the finding fires without naming the column. Phase 3 needed a raw schema_drift_snapshot() read and an offline per-column diff to classify three staging table findings. Recommendation: for the columns field, diff per column name (only-in-manifest / only-in-live / differing fields) and print those rows instead of the clipped arrays; keep the clip for other fields. Offline-testable against supabase/drift-manifest.json plus a mutated copy. | session 2026-08-18 Phase 3 repo-side codification (PR #2106) | 2026-08-18 | | #QSHHGK | P2 | rec | Nothing schedules a bundle-budget baseline refresh, so accumulated growth fails whichever unrelated PR lands last | The production baseline sat at ca788d41 (2026-08-13) untouched while main grew +8.03% by 2026-08-18, leaving ~2 points of headroom. PR #2096 (Dictionary) then failed Build at +10.5% for 2.7 points of its own weight. Re-baselined once in docs/evidence/bundle-budget-production-rebaseline-2026-08-18.md, but the same squeeze recurs unless a refresh has an owner or a trigger: options are a scheduled job that re-measures and opens a PR, a drift warning threshold below the failure threshold, or recording the baseline commit distance in the check output so staleness is visible before it blocks someone. | PR review of #2095/#2096, 2026-08-18 | 2026-08-18 | -| #5JK9FM | P3 | rec | PR template carries no RAG impact: guidance although pr-policy hard-blocks RAG-surface PRs without the line | .github/pull_request_template.md has zero occurrences of 'RAG impact', yet scripts/pr-policy.mjs ragImpactDeclared (lines 238-245) hard-blocks any PR touching a RAG-ranking-surface path unless the body carries a line matching 'RAG impact: ' with 'no ... behaviour change' or 'canary' and at least 12 characters. The authoring rule lives only in the pr-policy error string and AGENTS.md. Recommendation: add a commented placeholder line under ## Risk and rollout (or a dedicated ## RAG impact stanza) in the template with both canonical forms, so the exact-format contract is visible where the body is written; guard with the existing pr-policy self-test. | session 2026-08-18 Phase 3 repo-side codification (PR #2106) | 2026-08-18 | -| #SZGPAH | P2 | issue | tests/ui-tools-search-mode-mockup.spec.ts has two assertions stale on main, so the advisory lane is red for every UI PR | Verified 2026-08-18 on a clean origin/main worktree: node scripts/run-playwright.mjs --project=chromium-mockups tests/ui-tools-search-mode-mockup.spec.ts reports 2 failed \| 14 passed. The failures are 'desktop uses universal search and keeps results beside the selected-tool panel' (line 24) and 'phone filter sheet follows the shared local-filter behavior' (line 188, expecting the exact text '2 showing' in the filter sheet). Neither is caused by any open PR: PR #2095 was flagged for the second one while changing only mockup routes under src/app/mockups/caring-contacts. Advisory UI is non-blocking, so this stays red and trains reviewers to ignore the lane. Likely stale after the catalogue-toolbar work in #2086. Fix the assertions against current tools UI or quarantine per the flake ledger rules. | Copilot review triage on #2095/#2096, 2026-08-18 | 2026-08-18 | +| #SZGPAH | P2 | issue | tests/ui-tools-search-mode-mockup.spec.ts has two assertions stale on main, so the advisory lane is red for every UI PR | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): one of the two named assertions is provably gone. The phone test for the shared local-filter behaviour still exists but no longer asserts the exact text 2 showing anywhere in the file - that string has zero occurrences in tests/ui-tools-search-mode-mockup.spec.ts. The desktop test that keeps results beside the selected-tool panel still exists at line 24. NOT VERIFIED: whether the spec now passes - that needs node scripts/run-playwright.mjs --project=chromium-mockups on this file, which was not run. Re-measure before closing or quarantining. | Copilot review triage on #2095/#2096, 2026-08-18 | 2026-08-18 | | #6SMMB4 | P3 | task | Confirm D:\.npm-cache is a registered Dev Drive trusted cache, or Defender is scanning every npm ci | The repo lives on a Windows Dev Drive (D:, ReFS, 50 GB) and npm config get cache resolves to D:\.npm-cache, which is correctly on the same volume. Whether that path is registered as a Dev Drive TRUSTED cache is unverified: 'fsutil devdrv query D:' returns 'Failed to open the volume. Error 5: Access is denied' without elevation, and the non-elevated registry fallback (HKLM:\SYSTEM\CurrentControlSet\Control\FileSystem, FilterAttachModeOnDevDrive and DevDriveTrustSetting) reads empty. If it is not registered, Microsoft Defender real-time scanning runs over every npm ci — and this machine performs a lot of them: 21 D: worktrees each carry their own ~0.89 GB / 51,735-file node_modules, because npm extracts fresh copies rather than hardlinking from cache (ReFS does support hardlinks here, probed directly, but npm does not use them). Next: from an ELEVATED prompt run 'fsutil devdrv query D:' and, if the cache is not listed as trusted, 'fsutil devdrv trust D:\.npm-cache'. Cheap, one-off, no code change. Not blocking anything. | session 2026-08-18; fsutil Error 5 without elevation | 2026-08-18 | -| #TAQKCN | P3 | task | One hand-drawn SVG checkmark survived the Therapy Compass lucide sweep | recommend-screen.tsx:75 still inlines a 7-line hand-drawn tick inside the QUICK CONSTRAINTS pills; commit 7ce820f2c converted the other 32 glyphs in therapy-compass/** to lucide-react and deleted icons.tsx. It is a one-line swap to lucide Check, which the same file already imports for its copy buttons. Left out of the B3 Button conversion deliberately to keep that diff purely mechanical. Stop rule: it is cosmetic only - the pill's pressed state is already carried by aria-pressed plus border and text tokens, so nothing is colour-only while this waits. | session 2026-08-18 | 2026-08-18 | -| #E0N0QC | P3 | rec | scripts/probe-generation-quality.ts prints no answer text, so a blinded Gate E before/after read cannot come from probes | The probe reports structured generation_quality_gate_reasons only, by design, and never the answer text itself. That makes it impossible to produce a blinded before/after Gate E (answer-quality) read from probe output alone. Needs an owner-approved answer-text capture mode added to the probe, or a manual read of the app's rendered answer, before any future packet can close a blinded Gate E comparison this way. | scripts/probe-generation-quality.ts; #231 investigation, session 2026-08-18 | 2026-08-18 | -| #D6G8TC | P3 | task | Three Therapy Compass h1 elements still sit outside PageHeader, and the patient-sheet builder renders two of them at once | Stage D converted six therapy page headers to the shared PageHeader. Three h1 elements were deliberately left. (1) detail-screen.tsx:88 is the therapy record name inside the hero card, interleaved with an aria-live notice, aliases and a TagRow that PageHeader has no slots for - and PageHeader renders a
, which globals.css hides unconditionally under @media print, so adopting it would delete the record name from a printed therapy record. (2) sheets-screen.tsx:178 is the contentEditable title of the generated patient handout on --tc-paper-ink tokens; it is document content, not page chrome. (3) other-screen.tsx:28 is the centred placeholder hero, already at text-2xl. Consequence worth fixing: sheets-screen now renders the PageHeader h1 and the paper h1 simultaneously, so that page has two h1 elements - pre-existing, not introduced here, and the likely fix is demoting the paper title to h2 since the builder page is the document. Stop rule: do not solve (1) by adding a print exception for
; the element-name print rule is documented as transitional and must not be extended. | session 2026-08-18 | 2026-08-18 | | #GKFK9V | P3 | rec | Where the RAG improvement programme board lives, and to check it before starting RAG-surface work | The RAG improvement programme board lives at docs/rag-improvement/README.md (design), HANDOVER.md (packet status table + prompts) and COORDINATION.md (coordinator manual). Before starting any RAG-surface item, check that status table AND the open PR list (#292); packets S1-S3, G1, S4-S6 landed 2026-08-17/18; Track A complete. | docs/rag-improvement/README.md, HANDOVER.md, COORDINATION.md; session 2026-08-18 | 2026-08-18 | -| #RZQQBT | P3 | task | Confirm from its own log whether the PreCompact hook's output actually reaches model context | .claude/hooks/precompact-issues-capture.sh was added in PR #2113 to ask for /issues capture BEFORE compaction discards the follow-ups it wants recorded; the pre-existing issues-surface.sh reminder is a SessionStart hook and therefore fires AFTER compaction, when the material is already gone. Claude Code is known to inject hook stdout into model context for SessionStart, UserPromptSubmit, PreToolUse and PostToolUse. Whether it does so for PreCompact could NOT be determined: the installed CLI at %APPDATA%\npm\node_modules\@anthropic-ai\claude-code ships a compiled claude.exe with no inspectable JS bundle to grep. The hook therefore prints plain text rather than a hookSpecificOutput JSON envelope, so that if the platform does not inject it the operator still sees a clean transcript line instead of a raw JSON blob, and it appends one line per firing to a log so the question is answerable rather than permanently open. Next, after any compaction in a session using this repo: cat "$(git rev-parse --absolute-git-dir)/claude-precompact.log". Lines present but no reminder seen in context means the limit is real and the SessionStart backstop is carrying it. No lines at all means the registration is wrong. WARNING: that log lives under the worktree's own git dir, so it is destroyed when the worktree is removed — check it before cleaning up the worktree this was authored in. | PR #2113 .claude/hooks/precompact-issues-capture.sh | 2026-08-18 | -| #ZF006G | P3 | rec | Five independent SectionHeading implementations exist across modes with no shared recipe | clinical-dashboard/dashboard-shell.tsx:17, clinical-dashboard/search-pins-menu.tsx:97, formulation/formulation-ui.tsx:81, specifiers/specifier-ui.tsx:237, plus a fifth in therapy-compass/ui.tsx deleted unreferenced in 7d2c84a11. Same shape as the problem card-recipes.ts solved for cards, but spanning four modes. Two of the five already share an identical {eyebrow, title, body} signature - the natural shared contract. Stop rule: do not fold these into ui-primitives.tsx; COMPONENTS.md section 0.4 lists it as over-budget and slated to split. | session 2026-08-18 | 2026-08-18 | +| #RZQQBT | P3 | task | Confirm from its own log whether the PreCompact hook's output actually reaches model context | UPDATE 2026-08-21 (measured on this machine): still unanswered, and now with a concrete reason. The hook's own log has never been written - neither the main repository's .git/claude-precompact.log nor the per-worktree claude-precompact.log under the .git/worktrees admin directory exists. So no compaction has fired the hook in an inspectable session yet, and the question stays open rather than answered-negative. Next: after the next compaction in a session using this checkout, read that path; an entry proves the hook ran, and its absence in model context would then be the separate question. | PR #2113 .claude/hooks/precompact-issues-capture.sh | 2026-08-18 | | #CCZ4HB | P1 | rec | PR churn has exhausted the review-bot budget, so PRs are now landing with no automated review at all | CodeRabbit on PR #2113: '101 included PR reviews in the past 7 days; at that activity level, included reviews refill at 1 review per hour. Your organization has reached its usage spending cap.' The Codex connector reported its own usage limit on the same PR. Net effect: #2113 received ZERO automated review, and so will subsequent PRs until the cap resets or credits are added. AGENTS.md 'PR bundling' already measured the CI half of this cost on 2026-07-30 (437 PR-triggered runs over ~3 days, ~40% cancelled mid-run, ~12 Production-UI-hours burned on runs that never completed). This is the second bill for the same behaviour and the more dangerous one, because CI waste is money while missing review is undetected defects — and the PRs most likely to need review are the ones landing during a churn spike. The bundling rule exists as prose in AGENTS.md and is evidently not binding; the newtask skill also asks the question in prose. Decide whether it gets a gate. Note the repo has already learned this lesson once in a different area: .claude/hooks/pr-handoff-stop.sh states in its own header that 'prose rules in AGENTS.md have not held, a denied tool call does.' Next: decide between (a) a push/PR-creation gate that refuses a new branch when an open PR of the same scope exists, (b) raising the bot spending cap, or (c) accepting unreviewed merges deliberately rather than by accident. Stop rule: do not weaken any required check to compensate for missing bot review. | CodeRabbit + Codex connector comments on PR #2113, 2026-08-18; AGENTS.md 'PR bundling (reduce one-task-one-PR churn)' | 2026-08-18 | -| #EH9VA6 | P2 | rec | The serialized issues:reconcile operation has no interlock, and two reconcile PRs were open simultaneously on 2026-08-18 | AGENTS.md requires npm run issues:reconcile to run from ONE deliberately serialized fresh-base branch, but nothing enforces that: the reconciler's own guards (stale base, dirty canonical file, cross-worktree lock) are all local to a single machine and cannot see a second reconcile branch already pushed. On 2026-08-18 two were open at once — PR #2110/#2119 (claude/issues-reconcile-20260818, 'reconcile 74 queued requests') and PR #2120 (claude/issues-reconcile-2026-08-18-evening, 88 requests). Both merged. It only came out clean by luck of ordering: #2120 merged first and had already applied a superset of the 74, so GitHub's squash of #2119 collapsed to a 1-line review record and the canonical ledger was left correct (verified after the fact on main 4575cf57: 365 rows, 48 open, 0 pending, 339 applied, check:outstanding-issues and check:ledger-write-discipline both green). Had #2119 merged first, #2120's recorded transaction would no longer have equalled the canonical diff and check:ledger-write-discipline would have gone red on a branch that must never be synced from main — the documented recovery is to close the PR and redo the whole reconcile from a fresh base. Next: give the reconcile path a cheap pre-flight interlock rather than relying on operator discipline — e.g. have issues:reconcile refuse (or loudly warn) when git ls-remote --heads origin shows another unmerged branch carrying inbox renames under docs/outstanding-issues-inbox/applied/, which needs no provider tooling beyond a remote ref listing and no GitHub API call. Related to #292 (duplicate concurrent work) but distinct: this is a serialization invariant on a single canonical file, not two sessions building the same feature. | session 2026-08-18 evening reconcile; PRs #2119 and #2120 both open and merged same day | 2026-08-18 | -| #164Z0H | P3 | task | Confirm on a real Claude Code web session that the session-start hook now runs, after the exec-bit fix | PR #2113 fixed .claude/hooks/session-start.sh, which was checked in as mode 100644 while both sibling hooks were 100755, and was the only hook registered by bare path rather than through bash. What was PROVEN: the index mode, the bare-path registration, and that core.fileMode=false on this Windows ReFS Dev Drive hides both (a local chmod +x is a silent no-op; only git update-index --chmod=+x works). What was NOT proven: that it actually failed on a Linux web container, because no container was available to test from. The 100755/100755/100644 asymmetry makes accident overwhelmingly likely rather than a deliberate choice, and the script's whole body is gated on CLAUDE_CODE_REMOTE=true so the web container is the only place it does any work — it provisions the Node 24 the engine floor requires, after npm ci EBADENGINE blocked PRs #1611, #1697, #1705 and #1740. This is confirmation, not risk: the registration now uses bash "$CLAUDE_PROJECT_DIR/...", which removes the dependency on the mode entirely, and tests/session-start-hook.test.ts pins every hook at 100755 with LF-only line endings while tests/claude-code-settings.test.ts pins every hook command to start with an interpreter. Next: on the first Claude Code web session on this repo, check the session start output for the '[session-start] Using node ...' line and confirm npm ci ran. If it did not, the failure is something other than the exec bit and this item becomes a real defect rather than a confirmation. | PR #2113; AGENTS.md 'Claude Code hook scripts' | 2026-08-18 | -| #6GW95D | P2 | task | Nine landed worktrees are still on disk holding ~4.5 GB on a 51%-full Dev Drive; removal was deferred because the fleet was live | clean-worktree.mjs gained list-only --merged and --squashed in PR #2113 and identified 9 landed worktrees across the 50-worktree fleet, 5 of them on D:. Removal was NOT performed and must not be run blind. Re-verifying each candidate immediately before deletion, twice, showed the fleet is actively worked: database-coordination-chat-9c8cbd and database-drift-remeasure-phase2-7c4215 each held 2 unmerged commits despite the scan minutes earlier reporting '0 commits ahead', their newest files were written the same afternoon, and bundle-baseline had been switched to a different branch mid-scan and was running Playwright (the push guard named it as holding the heavy-run lease). Deleting any of them would have destroyed unmerged work. Next: run 'node scripts/clean-worktree.mjs --merged --squashed' when no other Codex/Gemini/Claude session is active, read the confidence line on each candidate, and re-run with --remove. Skip any candidate marked 'NOT fully corroborated' — that label means the patch-id test inferred the landing but some changed files still differ from origin/main, which is usually base churn but is not proof. D: was 25.3 GB of 50 GB used with roughly 19 GB of that duplicated node_modules across 21 worktrees at ~0.89 GB each. Ignore the C: worktrees entirely; they belong to Codex and Antigravity sessions. Stop rule: never pass --force to git worktree remove, and never remove a worktree that is ahead of origin/main. | PR #2113 scripts/clean-worktree.mjs; live fleet re-verification 2026-08-18 | 2026-08-18 | +| #164Z0H | P3 | task | Confirm on a real Claude Code web session that the session-start hook now runs, after the exec-bit fix | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): the repo-side half is now proven green. git ls-files -s .claude/hooks/ reports mode 100755 for all five hooks, session-start.sh included, so the 100644 asymmetry recorded here is fixed in the index rather than only on disk. What remains is exactly what this row was filed for and cannot be proven from Windows: one real Claude Code web/container session confirming the hook actually runs and provisions Node 24. | PR #2113; AGENTS.md 'Claude Code hook scripts' | 2026-08-18 | +| #6GW95D | P2 | task | Nine landed worktrees are still on disk holding ~4.5 GB on a 51%-full Dev Drive; removal was deferred because the fleet was live | UPDATE 2026-08-21 (measured on this machine): the figure recorded here is badly understated and moving the wrong way. git worktree list reported 69 registered worktrees under .claude/worktrees, not nine. At the ~0.89 GB / 51,735-file node_modules per worktree measured in #6SMMB4, even partial installs put this well beyond the 4.5 GB recorded, on a Dev Drive noted as 51% full. SAFETY, learned the hard way on 2026-08-21: a sweep removed an in-use worktree during this very session and destroyed its uncommitted work - see the separate row filed for that. Re-measure actual on-disk size, then prune, but confirm each worktree is landed, holds no uncommitted or unpushed work, AND is not currently checked out by a live agent session before removing it. | PR #2113 scripts/clean-worktree.mjs; live fleet re-verification 2026-08-18 | 2026-08-18 | | #VKH7N1 | P3 | rec | eval-canary neuroleptic-side-effect-escalation exceeded its 20 s latency SLO once | Run 32111839806 (canary pair 32100681177 -> 32111839806, otherwise green): strong generation took 20.2 s on neuroleptic-side-effect-escalation, flagged as a non-blocking latency advisory. Answer was still grounded via the source-backed extractive fallback. Watch on subsequent canaries; escalate only if it repeats or worsens. | docs/rag-improvement/HANDOVER.md packet table row S2, canary pair 32100681177 -> 32111839806 | 2026-08-18 | | #1K6T35 | P2 | issue | Point-in-time recovery is OFF on the live Supabase project, so the remediation plan's 'restore point before any mutating phase' rule cannot be met | Measured 2026-08-19 during the Phase 4 production window (forensics section 'Phase 4 completion', step 1). `supabase backups list --project-ref sjrfecxgysukkwxsowpy` returns pitr_enabled false, walg_enabled true, and seven retained daily physical backups, most recent COMPLETED 2026-08-17T20:33:28Z -- roughly 38 hours before that window opened. CONSEQUENCE: the recovery-point objective for the live clinical corpus (2851 documents, 70120 chunks) is up to ~24 hours, and the database-remediation plan's standing rule 'PITR/backup restore point captured before any mutating phase' CANNOT be satisfied on this project as currently configured. Phase 4 proceeded anyway and said so: every statement in it was index-only with an exact one-statement inverse (CREATE INDEX CONCURRENTLY <-> DROP INDEX CONCURRENTLY) and no data-loss surface, the same reasoning the 2026-08-14 incident window recorded. That reasoning does NOT generalise. Any future window that mutates DATA rather than indexes -- #022's BMJ attestation hosted apply, #036's public_corpus marker migration, #191's ACL consolidation, #057's restore/soak drill, or any reindex or backfill -- must not proceed on this precedent while the only restore point is a day-old physical backup. DECISION REQUIRED FROM THE OWNER, and it is dashboard plus billing work no agent can do: either enable PITR on the Supabase project (Database settings -> Add-ons -> Point in Time Recovery; it is a paid add-on, so this is a cost decision as well as a safety one), or deliberately accept the ~24h RPO and amend the plan's standing rule to say so, so that future sessions stop treating an unsatisfiable checklist item as if it had been met. Whichever is chosen, record it, because the current state is that the rule reads as satisfied by default when nobody checks. Consider re-grading this row to P1 if a ~24-hour worst-case data loss on the clinical corpus is judged unacceptable; it is filed P2 because the exposure is a deliberate platform configuration rather than a regression, and no data loss has occurred. Related: #057 (restore/rollback drill -- its value is limited while PITR is off), #188 and #196-#200 (DR codification). | Phase 4 production window 2026-08-19; docs/audit/live-drift-forensics-2026-08.md 'Phase 4 completion' step 1; PR #2151 | 2026-08-18 | -| #TF6TPJ | P2 | issue | Repeated main-merges on open PR branches cancel required CI, so 'PR required' reads red with zero failing jobs | PR #2143 went red on five consecutive heads and was merged past it. Job logs for runs 32170524256 (first) and 32178668323 (last) both report 'CANCELLED with no failing job' and zero failures: the first cancelled only lighthouse-budget while UI_FAST_RESULT and UI_RESULT were success; the last cancelled static-pr, coverage, production-ui-critical, production-ui and lighthouse-budget. Cause is cancel-in-progress firing on each new head, and something merged origin/main into that branch roughly every five minutes during a single working session, minting a head each time. #095 deliberately keeps a cancelled required job RED, which is correct, so the defect is the churn producing the heads, not the aggregate. Consequence beyond noise: the red is indistinguishable at a glance from a real break, and this PR was merged with it red. Next step is to find what is auto-syncing open PR branches (AGENTS.md 'Open PR branch sync (anti-churn)' says sync:pr-branches:apply is operator-run and should be late and once) and make it stop re-syncing a branch whose required CI is in flight. Related to #CCZ4HB, which records the same churn exhausting the review-bot budget. | PR #2143, runs 32170524256 and 32178668323, 2026-08-18 | 2026-08-18 | +| #TF6TPJ | P2 | issue | Repeated main-merges on open PR branches cancel required CI, so 'PR required' reads red with zero failing jobs | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): the guard that closes the shared root cause has landed. scripts/guard-push.mjs now carries an explicit Guard 2 in-flight CI push guard block naming #HSSHRG, with inFlightCiVerdict() and findInFlightCiRuns(), covering merge-main syncs made outside sync:pr-branches. #HSSHRG was closed on that evidence. This row should be re-checked against that guard and closed too if the false-red symptom is gone; it was NOT verified against a live PR from here, which needs GitHub access. | PR #2143, runs 32170524256 and 32178668323, 2026-08-18 | 2026-08-18 | | #P5542X | P2 | issue | pr-policy classifies the switch controlling whether unreviewed clinical content reaches production as clinicalRisk false, so no governance preflight is enforced | Found 2026-08-18 on PRs #2145 and #2150, both of which removed the Therapy production gate. classifyPullRequestFiles in scripts/pr-policy.mjs returned clinicalRisk false for a diff touching src/lib/app-modes.ts and src/lib/therapies.ts - the exact two modules deciding whether 205 clinically-unreviewed therapy records are reachable by users in production. Its clinicalRiskPatterns match src/lib/ only when the filename contains auth, permission, privacy, security, rag, retriev, rank, search, answer, clinical, citation, source, document, upload or download; app-modes and therapies match none. The data patterns (src/data, data, public/therapy-compass-data) match the records themselves but not the code gating their reachability. Both PRs completed a governance preflight voluntarily, which is exactly the fragility: the next such change may not. Consider matching on reachability/exposure surfaces, or treating any diff that changes an app-mode devOnly flag or a review-status filter as clinical-risk. Stop rule: do not widen the patterns so far that ordinary UI work trips the preflight - the classifier comment already warns that presentation files are not clinical-risk merely for living under a clinically-named directory, and that judgement is correct. | session 2026-08-18 | 2026-08-18 | -| #VTEW3W | P3 | rec | therapyBtn still dresses 30 raw controls across 8 therapy files with no shared equivalent | Measured 2026-08-18 on main at adf93a75. src/components/therapy-compass/controls.ts exports therapyBtn, used at 30 call sites across therapy-card.tsx (2), detail-screen (2), pathways-screen (4), compare-screen (3), search-screen (2), sheets-screen (8), recommend-screen (6) and brief-screen (3). PR #2122 deleted the three button recipes it sat beside (commandControl, outlineControl, iconControl) and converted their 19 call sites to the shared Button, but deliberately kept therapyBtn: its remaining consumers are list rows, disclosure headers, section toggles and chips - controls that are not Buttons in the design-system sense, so forcing them onto the Button variants would have been wrong. The gap is that nothing shared covers them either, so therapy carries its own focus ring, hover lift and dual disabled encoding (native disabled plus aria-disabled) while every other mode hand-rolls or omits equivalents. Worth deciding whether the design system should own an interactive-row or quiet-control recipe, which would serve more than therapy - the same question card-recipes.ts answered for cards. Note therapyBtn already composes the shared focusRing from card-recipes.ts, so the focus contract is not forked; what is local is the hover/press motion and the disabled encoding. Stop rule: do not resolve this by converting the 30 sites to Button variants - that was considered and rejected during #2122 because a list row is not a button, and it would put button chrome on card-like surfaces. | session 2026-08-18 | 2026-08-18 | +| #VTEW3W | P3 | rec | therapyBtn still dresses 30 raw controls across 8 therapy files with no shared equivalent | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): the count has fallen. therapyBtn now appears 21 times across 8 files under src/components/therapy-compass/, against the 30 call sites measured 2026-08-18 at adf93a75. The design question is unchanged - nothing shared covers list rows, disclosure headers, section toggles and chips - and the stop rule still stands: do not resolve this by converting the remaining sites to Button variants. | session 2026-08-18 | 2026-08-18 | | #GBBYTA | P2 | task | Hoist the filtered-zero empty state out of the documents results grid so it can sit flush under the results band | SearchResultsEmptyState renders the filtered-to-zero case nested two divs inside the results grid in document-search-results.tsx (grid gap-3 wrapper, with a conditional 'N results after filters' pill above it), not adjacent to SearchResultsHeaderBand. That nesting blocks the inline treatment evaluated for PR #2147: recovery cannot share the band's bottom edge, the band's data-tone lead is not adjacent so the panel has no state mark, and the filter chips are not near enough for the panel to point at them instead of duplicating them. PR #2147 therefore shipped the self-contained rail panel, which does not depend on adjacency. Hoisting the state to a sibling of the band would let the panel drop its own rail, its named-removal row and its eyebrow — roughly a third less height on phones. Only the documents nesting was verified; the other five consumers (favourites, calculators, forms, services, therapy-compass) were not checked and may nest the same way. | PR #2147 design review; document-search-results.tsx:1558 | 2026-08-18 | | #0HFDWD | P2 | issue | CI change-scope reports UI_CHANGED false for changes that alter which modes render, so Production UI is skipped on user-facing work | Found 2026-08-18 on PR #2145. The CI run recorded UI_CHANGED false, UI_RESULT skipped and UI_FAST_RESULT skipped for a diff that removed devOnly from the therapy-compass mode and switched off the production record filter - a change that alters which modes appear in the shell for every user. Production UI therefore never ran. The uiPatterns in scripts/pr-policy.mjs and the equivalent scope detection in scripts/ci-change-scope.mjs match src/app/ (non-api), src/components/, src/styles/, public/, tests/ui-*.spec.ts and playwright config; the diff touched only src/lib/app-modes.ts, src/lib/therapies.ts and unit tests, so nothing matched. PR #2150 supersedes that change and adds a visible component, and would still not trip the classifier for its src/lib half. Browser coverage for both was supplied only by a local verify:ui run (447 passed), which no policy required. Consider treating src/lib/app-modes.ts as UI scope, since it is the mode registry the shell renders from, and auditing which other src/lib modules feed rendering. Stop rule: do not make all of src/lib UI scope - that would run a 20-minute Chromium gate on every library change and reintroduce the cancellation waste documented in docs/testing.md. | session 2026-08-18 | 2026-08-18 | | #EP1BQS | P3 | rec | The RAG improvement programme carries its own status table with no session-start link back to the ledger | NARROWED 2026-08-18 before filing: an earlier read of this found zero ledger references to docs/rag-improvement/. That is no longer true — #231 now cites it (5 mentions), so the ledger-to-programme direction exists. What remains is the reverse direction and ownership. docs/rag-improvement/HANDOVER.md instructs every implementing session to read its packet, then README.md, then docs/rag-behaviour/, and to update a per-session status table (section 2) — it does not tell that session to read docs/outstanding-issues.md or to check open PRs for the surface its packet touches. So a session entering through HANDOVER.md never sees the canonical queue, and no row owns the programme itself. AGENTS.md designates docs/outstanding-issues.md as the single universal cross-session ledger and warns that detailed runbooks must not become a second status ledger; a per-session status table is close to that line. This is the documented precondition for #292 (two assistants shipped the same queued conversion four hours apart, one PR closed as a duplicate) and #301 (two sessions built #262 part 3 in parallel), and it is not hypothetical for this programme: the same check-before-filing discipline is what stopped this very row from being filed with a stale premise. Next (cheap, docs-only, no behaviour change): add to HANDOVER.md section 4's session-start checklist a step to read the recommended queue and check open PRs for the packet's surface, and decide whether the programme warrants one owning ledger row so its state is visible from the canonical side. Stop: do not duplicate the programme's tracks as ledger rows — that would create the second queue this row exists to prevent, in the other direction. Design authority stays with docs/rag-improvement/README.md and protected-surface rules with docs/rag-behaviour/. | session 2026-08-12/08-18 cross-session review; verified against origin/main 622f1fb | 2026-08-18 | @@ -140,21 +130,21 @@ removed after current-main verification; it is not missing recommended work. | #NEBJAM | P2 | rec | Therapy keeps a private eight-component UI kit whose shared equivalents it imports zero times, including the badge that carries review status | Measured 2026-08-18 on main at adf93a75, after PRs #2122 and #2150. src/components/therapy-compass/ui.tsx exports Tag, TagRow, StatusBadge, IconTile, LoadingState, EmptyState, Eyebrow and Meter. Shared equivalents exist and Therapy imports them zero times: ui/chip.tsx for Tag/TagRow (Chip is imported in exactly one therapy file), ui/status-mark.tsx for StatusBadge, ui/error-state.tsx for EmptyState, ui/progress.tsx for Meter, the eyebrowText primitive for Eyebrow, and category-icon-tile.tsx for IconTile. This is the same class of duplication card-recipes.ts was written to end for cards, one layer down, and it survived the #2122 convergence because that work targeted Button, card surfaces and page headers only. Highest-consequence piece: StatusBadge renders the Needs source review label, which since #2150 is the per-record half of the only protection standing between an unreviewed therapy record and a clinical decision - and it is a module-private implementation no shared contract governs. Two migration hazards to respect. (1) StatusBadge pairs its warning tone with a TriangleAlert glyph via reviewStatusMeta in data/select.ts; any swap must preserve that shape channel or the warning becomes colour-only and trips the status-colour boundary ratchet. (2) Meter's colour-only status was deliberately fixed in commit 8c791a1ad by naming the completeness band in text; a naive swap to shared Progress would regress it. Suggested order: Eyebrow and Tag/TagRow first (lowest risk, no status semantics), then EmptyState and IconTile, then Meter and StatusBadge last with the colour-only contract tests extended first. Stop rule: do not fold any of these into ui-primitives.tsx; COMPONENTS.md section 0.4 lists that module as over-budget and slated to split. Related: #ZF006G tracks the parallel SectionHeading duplication across four modes. | session 2026-08-18 | 2026-08-18 | | #SBKXZ7 | P2 | task | Therapy sign-off has no tooling: nothing stops reviewStatus reviewed being set with an empty checklist, and there is no reviewer attribution | Re-queued 2026-08-18 after PR #2145 was closed in favour of #2150, which supersedes the exposure change but does not carry this follow-up. Therapy now ships in production with its review state disclosed rather than hidden, so sign-off is the remaining clinical work. Three gaps. (1) reviewStatus is a bare string in src/data/therapies-source.json; a record can be flipped to reviewed with all seven reviewChecklist booleans still false and nothing detects it. Needs a script that refuses the flip unless the checklist is complete, plus a contract test pinning reviewed implies full checklist. (2) No attribution: none of the 44 record fields carries reviewedBy or reviewedAt, so a sign-off cannot record who signed or when - the same defect #318 flags against the medication interaction lexicon. (3) No review workflow: 205 records x 7 checks is 1435 clinical judgements by hand; a CLI that walks records, shows the fields each check covers, and writes the decision with attribution would make it tractable. State at re-queue: 205 records, all reviewStatus needs_review, all seven checklist booleans false, reviewCompleteness 57-71 with zero records complete. The catalogue notice #2150 adds reads from THERAPY_CATALOGUE_SUMMARY.needsReviewCount and disappears when that reaches zero, so completing sign-off is what retires it. Stop rule: an assistant must never tick clinicalAccuracyReviewed, sourceChecked, evidenceAppraised, safetyCautionsChecked or patientExplanationChecked - those are qualified-clinician attestations. proofread and australianEnglishChecked are non-clinical and may be done with attribution. | session 2026-08-18 | 2026-08-18 | | #BSBE9B | P3 | task | Docling lab fixtures.v2 table-hardness corpus: add unruled, merged-cell, and rotated-header tables so the Gate B table-heavy improvement leg has measurable headroom | The v1 corpus's table strata are cleanly ruled grids on which the legacy extractor already scores cell F1 1.0 (S6 smoke run and the S6b Gate B run), so the pre-agreed table-heavy improvement target was set to 0 pp (parity at ceiling) by owner decision on 2026-08-18. Before any table-heavy delta is treated as decisive for a Docling promotion beyond B4 shadow design, add a docling-lab-fixtures.v2 stratum set where the legacy find_tables path is expected to degrade: unruled tables, merged cells, rotated headers. Fixture-hardness change only — eval/docling/ manifest + generator, no worker or extractor edit. See eval/docling/README.md 'Known limitation (v1 corpus)' and docs/rag-improvement/gate-b-decision-record-2026-08-18.md. | packet S6b (Gate B run), owner threshold decision 2026-08-18 | 2026-08-18 | -| #HSSHRG | P2 | issue | The in-flight-CI guard closed as #145 does not cover merge-main syncs made outside sync:pr-branches — PR #2149 lost 8 of 9 CI cycles to self-inflicted cancellation | #145 'Branch syncs could repeatedly cancel healthy in-flight CI' is recorded RESOLVED 2026-07-30 on the basis that the operator helper queries Actions for the current head and skips update-branch while the required workflow is queued or in progress. That guard lives in scripts/sync-pr-branches; it does not constrain a plain 'git merge origin/main' + push, nor GitHub's Update branch button, and PR #2149 shows the uncovered path is the one actually used. Between 18:51 and 20:16 on 2026-08-18, PR #2149 received SEVEN 'Merge branch main into claude/diagnostic-criteria-duplication-udg99e' commits (18:51, 19:03, 19:18, 19:37, 19:42, 19:54, 20:16) and two more by 21:37, each a new head that restarted CI and killed the run in flight via cancel-in-progress. Across nine cycles the aggregate went red eight times as 'CANCELLED with no failing job' and never once for a genuine failure of the diff. The cancellations widened as the interval tightened: cycle 2 lost only lighthouse-budget, cycle 3 only production-ui, cycle 4 lost coverage + production-ui-critical + production-ui + lighthouse-budget, and the 20:18 cycle lost all ten jobs including 'changes' itself, whose *_CHANGED outputs were empty because scope detection never completed. Real cost: the one genuine signal in 3.5 hours — a Production UI failure in tests/ui-forms-section-nav.spec.ts — had its re-run cancelled and its next cycle cancelled, so the PR merged with that question permanently unanswered (see sibling record). Note #095 behaved exactly as designed throughout: the aggregate correctly distinguished cancellation from failure and correctly stayed red, since a cancelled job verifies nothing. The defect is upstream of the aggregate. Next: decide whether the in-flight check should move from the operator helper to a place every sync path passes through — candidate options are (a) extend .githooks/pre-push / scripts/guard-push.mjs to warn or block when pushing a merge-only commit to a PR branch whose required CI is queued or in progress, (b) document the Update-branch button as prohibited for any PR with CI in flight rather than only for ledger-touching PRs as AGENTS.md currently states, or (c) accept it and reduce sync frequency. Stop rule: do not disable cancel-in-progress for pull requests to fix this — it is deliberate for PRs, is pinned by tests/ci-cache-safety.test.ts, and its base-branch exemption exists for a separate reason. | PR #2149 commit list and runs 32173376307 / 32173687350 / 32174490902 / 32181371909 / 32185492075 / 32188061671; session 2026-08-18 | 2026-08-18 | | #5DYBQQ | P2 | issue | tests/ui-forms-section-nav.spec.ts 'expands information previews into one continuous answer' failed on PR #2149 and was never reproduced — CI's path-scoped UI jobs can hide a forms regression on main | On PR #2149 (a DSM-only, two-file change) Production UI shard 3 failed 166 passed / 1 failed on tests/ui-forms-section-nav.spec.ts:53 'expands information previews into one continuous answer'. The assertion is 'await expect(trigger.getByText(preview)).toHaveCount(0)' at line 65: after clicking the 'Does not authorise' trigger on the form detail route, the preview text 'Psychiatric treatment or detention beyond the linked authority.' must leave the trigger and appear in the panel. It stayed in the trigger. NOT a timing race: the locator polled 24 times over the full 10s timeout and resolved to 1 element every time, so it was a stable wrong state, not a slow transition. NOT attributable to #2149: that PR changed exactly src/components/dsm/dsm-diagnosis-page.tsx and one ledger record, zero forms files; the behaviour lives in src/components/forms/form-detail-page.tsx which shares nothing with the DSM page. NOT a known flake: the identity is absent from tests/flake-ledger.json. NOT explained by recent forms history: the only recent main commit touching forms is 495e097 (#2139), which edited forms-home-page.tsx only (caveat-footer removal), a different component from the form detail page under test. Reproduction was never obtained: the job was re-run once and that re-run was cancelled, and the following full cycle was cancelled too, both by the merge-main sync loop described in the sibling record. #2149 then merged with the question open. The reason this can hide: CI UI jobs are path-scoped, so a docs-only push to main SKIPS Production UI entirely — verified in main run 32183858120 where 'Production UI', 'Production UI critical', 'Build' and 'Lighthouse budget' all report conclusion 'skipped'. The repo's own CI-triage bot nevertheless cited that run as 'Compared with main CI run #12334 (success)', which is an aggregate-level comparison that never exercised this test and must not be read as a green baseline. Given main's docs-heavy traffic, a genuine forms regression could sit on main unexercised. Next: run the single spec against current main to settle it — 'npx playwright test tests/ui-forms-section-nav.spec.ts --project=chromium' after 'npm run ensure' — then either fix form-detail-page.tsx so the preview moves out of the trigger on expand, or, if it passes repeatedly, add the identity to tests/flake-ledger.json under the documented three-reproductions-on-one-SHA rule. Stop rule: do not quarantine on a single observation, and do not weaken the assertion to green it. Note this could not be reproduced locally in the Claude web container because Playwright pins chromium revision 1234 while the image ships 1194 (see #255); forcing a mismatched browser path is disallowed. | PR #2149 run 32185492075 job 95868823944 (Production UI (3)); main comparison run 32183858120; session 2026-08-18 | 2026-08-18 | | #D8JBCV | P2 | issue | /tools on a phone is the only mode home with no visible patient-identifiable-information warning | Tools is the sole route setting mobileHomeComposerPlacement: 'footer' (src/lib/search-shell-props.ts). showsComposerPrivacyNotice in master-search-header.tsx:1813 is 'usesPhoneSearchLayout ? isDesktopHomeComposer : true', so the phone footer dock suppresses both the 'Do not enter patient-identifiable information.' line and the Privacy and data processing link. The composer placement matches the documented exception in docs/search-chrome-behaviour.md row 2, but the docs do not record that the exception costs the governance copy. Needs an owner decision for a clinical product. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | -| #97VQK5 | P3 | rec | Mode home copy drift: three placeholder-punctuation conventions, inconsistent heading levels, and a stale docs/site-map.md mode index | Placeholders use ASCII '...' (answer, documents, services, forms, favourites, dsm, specifiers, formulation, prescribing, tools, calculators), Unicode '…' (therapy-compass, factsheets, dictionary), and no terminator at all on differentials ('Ask or search a presentation'). Heading level is h2 on answer/documents/prescribing and h1 elsewhere, so the Documents home has an h2 and no h1 anywhere on the page. docs/site-map.md's mode page index covers 11 modes — Therapy, Factsheets, Dictionary and Calculators have no row — and CLAUDE.md still says '13 app modes' when app-modes.ts defines 15. Pick one convention per axis and refresh the generated docs. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | -| #6K9YGQ | P2 | issue | Three standalone mode homes have no loading.tsx (/medications, /calculators, /dictionary) and the contract test cannot catch it | All three are in standaloneModeHomePaths (src/lib/search-route-ownership.ts) and chrome invariant 18 requires ModeHomeRouteLoading, but tests/mode-home-loading-contract.test.ts enumerates only ten routes and omits these three, so the gap is invisible to CI. /documents is legitimately exempt (dashboard-owned). Fix: add the three loading.tsx files and derive the contract test's route list from standaloneModeHomePaths instead of a hand-written list. Verified by direct filesystem check during the PR #2160 audit. | PR #2160 cross-mode home audit | 2026-08-18 | +| #97VQK5 | P3 | rec | Mode home copy drift: three placeholder-punctuation conventions, inconsistent heading levels, and a stale docs/site-map.md mode index | UPDATE 2026-08-21 (repo read on main at 1cc0d2987): the documentation half is DONE. CLAUDE.md now reads 15 app modes, and the mode index in docs/site-map.md now carries rows for Calculators, Therapy, Factsheets and Dictionary, so the four missing modes and the stale count are both corrected. REMAINING SCOPE, and the only reason this row stays open: the placeholder-punctuation split (ASCII three dots vs Unicode ellipsis vs no terminator on differentials) and the heading-level inconsistency (h2 on answer/documents/prescribing, h1 elsewhere, so the Documents home has no h1). Both are decisions, and the heading change moves visual baselines. | PR #2160 cross-mode home audit | 2026-08-18 | | #YJ3R7Y | P3 | issue | Tools and Favourites bespoke home composer slots skip the SSR height reservation chrome invariant 15 requires | ModeHomeTemplate renders its composer slot with data-composer-reserve='pending' plus min-h tokens (mode-home-template.tsx:316-320) so the hero does not shift when the portal attaches. The two bespoke homes hand-roll the slot without either: favourites-command-library-page.tsx:1426 and tools-search-results-page.tsx:353, plus favourites-hub.tsx:187. Those three get no SSR height reservation, which is the CLS that invariant 15 exists to prevent. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | | #TWKWE4 | P2 | issue | Mode homes: two competing title systems disagree for 8 of 13 modes (sharedHomePresentation vs hard-coded standalone titles) | src/lib/ui-copy.ts sharedHomePresentation drives the shared home /, while each standalone *-home-page.tsx hard-codes its own title. Its doc comment claims each entry mirrors the standalone home 'so a clinician sees the same words whichever door they came through' — untrue today: Documents/Clinical Documents, Services/Clinical Services, Forms/Clinical Forms, Differentials/Differential Diagnosis, Specifiers/Diagnostic Specifiers, Formulation/Clinical Formulation, Medication/Medication Guidance, Therapy/Therapy Compass. Either derive one list from the other or correct the comment. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | -| #0EKBGC | P3 | issue | Three mode homes override the canonical APP_MODE_ICON glyph (services, forms, dictionary) | services uses Users (canonical route), forms uses FileText (canonical fileSignature — and identical to the Documents home tile), dictionary uses BookOpen (canonical bookMarked). Same class as the therapy-compass magnifier fixed in PR #2160, which now derives from appModeIcons. Each remaining one needs its own visual-baseline re-adoption, so they were left out of that PR. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | | #V0EDR4 | P3 | issue | /favourites and /?mode=favourites render visibly different homes for the same mode | The standalone hero lockup was deliberately deleted from favourites-command-library-page.tsx (ledger #164), but the dashboard variant FavouritesHub (src/components/clinical-dashboard/favourites-hub.tsx:179) still renders ModeHomeHero with 'Favourites / Saved notes, sources, and sets.' So the same mode looks different depending on the door. Decide which treatment is canonical and apply it to both. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | | #90Y0FD | P3 | rec | Mode home suggestion data is duplicated across three unrelated sources | searchCommandSurfaceByMode examples/suggestions (src/lib/search-command-surface.ts) drive the Try this ticket, rotating hint and prompt chips; per-page pills arrays (e.g. therapy-compass/screens/home-screen.tsx:14, services-home-page.tsx) drive the mode-home pill row; src/lib/tools-catalog.ts:348 is a third. Only the first drives the ticket, so after PR #2160 the Therapy home advertises two different suggestion sets — its five pills and the three ticket examples. Reconcile to one source per mode. Found during the PR #2160 cross-mode audit. | PR #2160 cross-mode home audit | 2026-08-18 | | #JVYQEM | P2 | issue | Mode-home composer reserve does not account for the suggestion ticket, so every ticket-bearing home carries a ~0.035 CLS shift | ModeHomeTemplate reserves the composer slot with --spacing-mode-home-composer-phone (6.625rem) / --spacing-mode-home-composer-wide (5.5rem), but the portal content is UniversalSearchCommandSurface, which renders SmartRotatingHint (phone ticket) and the sm+ rotating line/prompt-chip row ABOVE the composer inside that same slot. The reserve therefore under-accounts, and the portal attaching post-hydration shifts content — the defect class chrome invariant 15 exists to prevent. Evidence from the PR #2160 Lighthouse run: mobile-dsm baseline CLS 0.0353, mobile-forms 0.088, mobile-root 0.016, while mobile-therapy-compass was 0.000 purely because Therapy had no command-surface entry and so rendered no ticket. Restoring the ticket moved Therapy to 0.032, matching its peers. Fix: raise the reserve tokens to include the hint row height (or reserve it separately), which should take every mode home toward ~0. Touches all 15 mode homes, so it needs verify:phone-chrome plus Lighthouse and visual baseline re-adoption — deliberately not bundled into PR #2160. | PR #2160 Lighthouse budget failure | 2026-08-18 | | #HVTYAT | P2 | issue | OpenAI zero-data-retention status contradicts itself: cross-border doc says no, ledger #053 says verified | docs/openai-cross-border-basis.md §8 (dated 2026-07-14) records ZDR 'no', DPA executed 'no', Australia data residency 'not enabled'. Completed ledger row #053 (2026-08-18) states the cross-border package was executed and 'verified OpenAI data controls with input/output data sharing disabled and API zero data retention'. One of the two is wrong. This blocks the /privacy page from telling clinicians what actually happens to question text at the provider: the page currently states only code-verifiable request controls (store:false, no raw owner identifier, requested prompt-cache lifetime) and deliberately makes no ZDR or no-training claim, which is correct under either reading but weaker than it could be. Next step: an operator confirms the live OpenAI project's data controls, then either §8's status table is filled in and the page's External provider processing section is strengthened, or #053 is corrected. | src/lib/privacy-page-content.tsx, docs/openai-cross-border-basis.md §8, docs/privacy-impact-assessment.md PIA-1/PIA-6 | 2026-08-19 | | #1VFSYF | P3 | task | Close the four operator unknowns the B4 shadow-extraction runbook could not answer from the repository: Railway variable-change behaviour, the shadow-record read path, the timeout rollback threshold, and the worker memory limit/peak | docs/worker-deploy-runbook.md section 3 (added 2026-08-20) is the Gate F runbook for packet B4 docling shadow extraction. Three values in it could not be verified from the repository and are flagged as such in the text. (1) ROLLBACK TIME TO EFFECT: the one-step rollback is WORKER_DOCUMENT_EXTRACTOR_MODE=legacy on the Railway worker service, but whether Railway restarts or fully rebuilds the service on a variable change is not recorded anywhere in this repo, so the runbook can only bound the worker-side cost (SIGTERM batch drain plus at most one 120 s docling window). One dashboard observation settles it; record the answer in section 3.7. (2) NO READ PATH: no npm script reads, exports, or clears documents.metadata.shadow_extraction, so the first-24-hours watch in section 3.6 is a hand-run SQL query in the Supabase editor. A small read-only aggregation script (outcome counts, wall_ms and peak_rss_bytes percentiles, delta ratios) would make the watch repeatable and reviewable instead of ad hoc. (3) TIMEOUT ROLLBACK THRESHOLD: section 3.7 trigger 5 proposes rolling back when more than 10 percent of cohort runs time out. That number is a proposed operating rule, not a measured value, and needs owner ratification or replacement once real-corpus wall_ms data exists. Also unverified: the memory headroom precondition in section 3.2 asks the operator to confirm the Railway worker memory limit and observed peak, neither of which is in-repo. Stop: no code, worker, or provider change is implied by this row; it is documentation completion plus one small offline read-only script. | docs/worker-deploy-runbook.md section 3 (Gate F runbook PR, 2026-08-21); PR #2170 squash 5437c309f | 2026-08-20 | -| #VZN8G3 | P3 | task | docling-lab-fixtures.v2: add unruled, merged-cell and rotated-header table fixtures to the Docling lab before any table-quality promotion argument | Gate B (PR #2154, run 32176604314 at 8a92378) passed with the table-heavy leg at parity-on-ceiling: the v1 fixture set under eval/docling/ cannot separate docling from the legacy extractor on tables because every fixture is ruled/simple. Packet B4 (PR #2170, squash 5437c309f) shipped shadow mode as measurements only. Before anyone argues docling table quality (a precondition for any promotion beyond shadow, README §B4), author docling-lab-fixtures.v2: unruled tables, merged/spanning cells, rotated headers, plus expected exact number/unit/comparator checks, and re-run the lab (docling-lab.yml dispatch = hosted CI, owner approval). Synthetic/public sources only; hashed lockfile unchanged; aggregate-only report. Stop: no worker, extractor, or database change; shadow numbers remain measurements until v2 says otherwise. | RAG programme coordinator, post-B4 (2026-08-19) | 2026-08-19 | | #M54C4N | P2 | issue | live-drift's Align migration history step fails on PGRST106 (supabase_migrations not exposed to PostgREST), so the job stays red and pinned issue #1963 cannot self-close even with zero drift | Exposed 2026-08-19 by Phase 6.2: live-drift run 32251326536 reported No unexpected schema drift (all 20 history rows allowed) but concluded failure because the next step, Align migration history for Supabase Preview (npm run check:migration-history, scripts/check-migration-history-alignment.ts, added in Phase 0 PR #1939), ran for the first time ever -- it was skipped on every earlier run because the compare step failed first, and the last green run (29700973962, 2026-07-19) predates it. It reads supabase_migrations.schema_migrations through PostgREST with Accept-Profile, which this project has never exposed (406 PGRST106: only public, graphql_public). The routing job therefore keeps #1963 open on job result even with an empty findings block. OPTIONS (owner decision): (a) expose supabase_migrations read-only to the service role in the dashboard; (b) rewrite the read onto the management API / supabase migration list using the SUPABASE_ACCESS_TOKEN secret (#183); (c) add a service-role RPC listing versions (new migration, own window). Until fixed, weekly live-drift stays red on that step alone; drift itself is green. Evidence: forensics section 6.2 completion step 6. | Phase 6.2 session 2026-08-19; live-drift run 32251326536 | 2026-08-19 | +| #1YPV51 | P1 | task | Reopen #318: the medication interaction lexicon clinical review and sign-off has not actually happened; the lexicon remains clinically unreviewed | #318 was closed 2026-08-18 with outcome 'Completed Clinical Lead medication interaction lexicon review and sign-off'. The repo owner confirmed in chat on 2026-08-21 that this review and sign-off was not performed by the Clinical Lead. The medication interaction lexicon therefore remains clinically unreviewed and its sign-off block is empty, exactly as the original #318 described before its (inaccurate) closure. This must be reopened and tracked through to a genuine Clinical Lead review and sign-off; do not close again without explicit owner confirmation that the review was actually carried out. Related: #SBKXZ7 flags the same missing-attribution gap (no reviewedBy/reviewedAt) for therapy sign-off. | session 2026-08-21 ledger reconciliation and docs-truth pass; owner confirmed in chat the 2026-08-18 closure was inaccurate | 2026-08-20 | +| #XCAX01 | P2 | issue | A worktree sweep deleted an in-use worktree mid-session on 2026-08-21 and destroyed its uncommitted work; nothing checks whether a worktree is live before removal | Reproduced by loss, not by test. On 2026-08-21 an agent session was working in .claude/worktrees/task-ledger-review-bee095 with 25 staged-but-uncommitted files. Mid-task the directory was emptied, its .git/worktrees admin directory (and therefore its index) was deleted, and the worktree was deregistered, while the session was still running. The staged blobs became unreachable and the work was lost; only the branch ref survived, still at its unchanged base 1cc0d298774e. Every other worktree on disk carried a modification timestamp inside the same twenty-minute span, so this was a sweep across the whole directory rather than a one-off. TWO CONSEQUENCES WORTH SEPARATING. (1) Data loss: removal considered neither uncommitted/untracked content nor whether a process was live in the directory. (2) Silent corruption of the surviving session: after the removal, git commands issued from the deleted path resolved upward to the main checkout D:/Repos/Database, which was on another agent's feature branch with uncommitted modifications - so an unlucky commit would have landed in a different session's branch and working tree. NEXT: before any automated worktree removal, require all three of (a) the branch is merged or its tip is pushed, (b) git status --porcelain --untracked-files=all is empty, and (c) no live process holds the directory; and skip rather than force when any check cannot be evaluated. Prefer a report-only default with an explicit apply flag, matching how sync:pr-branches already separates dry-run from apply. RELATED: #6GW95D records the disk pressure that motivates sweeping, and now carries the same safety caveat. | Session 2026-08-21; worktree task-ledger-review-bee095 removed while in use | 2026-08-20 | +| #S19JRT | P2 | task | Add the DB-side structural constraint backing the source_metadata pin, or document why the data-backed pin is sufficient | Re-files #343, closed 2026-08-18 with outcome 'Made retrieval row contract source_metadata schema structural and nullish' -- that outcome is false. Verified 2026-08-21: PR #2107 loosened the source_metadata pin in src/lib/rag/rag-row-contracts.ts to .nullish(); PR #2121 restored the strict .nullable()-required-key pin (git log: ce702ba68 then 4575cf57a). The comment at rag-row-contracts.ts:44-49 explicitly reads 'PR #2107 loosened it to .nullish() and this PR restores it. See docs/outstanding-issues.md #343 for the constraint-backing follow-up.' The DB-side structural constraint (check (jsonb_typeof(metadata) = 'object')) was never added: grep of supabase/schema.sql and supabase/migrations/ finds only 'metadata jsonb not null default {}::jsonb' with no jsonb_typeof check anywhere. The cancelled duplicate #ND10QT record itself states '#343, which is still open', confirming the two closures landed inconsistently. Actionable follow-up: add the check (jsonb_typeof(metadata) = 'object') constraint on documents.metadata with a fail-fast validation guard migration per AGENTS.md's guard-migration contract, or record in this row why the Zod-level pin in rag-row-contracts.ts is sufficient without a DB constraint. | session 2026-08-21 ledger reconciliation and docs-truth pass | 2026-08-20 | +| #3514B7 | P2 | issue | Live production carries migration 20260820120000_migration_history_versions_rpc, which exists nowhere in the repository | Found 2026-08-21 by a read-only Supabase MCP list_migrations against production ref sjrfecxgysukkwxsowpy (Clinical KB Database), compared against main at 1cc0d2987. Production's newest recorded version is 20260820120000 migration_history_versions_rpc. supabase/migrations/ holds 210 files and its newest is 20260819110500_validate_history_function_bodies.sql; there is no 20260820 file, and the string migration_history_versions has zero occurrences anywhere under supabase/, scripts/ or src/. The live database is therefore one migration ahead of the repository, and that migration's SQL is not under version control here. LIKELY RELATED, and the reason this matters rather than being cosmetic: pending request cc60253d (2026-08-19) records that live-drift's Align migration history step fails on PGRST106 because supabase_migrations is not exposed to PostgREST. An RPC named migration_history_versions is exactly the shape of a fix for that, applied live on 2026-08-20 with the repo side still unmerged. NOT VERIFIED from this session: whether an open PR carries the file, which needs GitHub access. NEXT: identify the PR or session that applied it; if none exists, capture the live function definition and land it as a forward migration so schema.sql, the drift manifest and the migration chain agree. STOP: do not re-apply, repair, or mark-apply anything on production to resolve this - it is a repo-side reconciliation, and any live mutation needs its own approved window under the guard-migration contract. | Read-only Supabase MCP list_migrations on sjrfecxgysukkwxsowpy, 2026-08-21; compared to supabase/migrations on main 1cc0d2987; relates to pending request cc60253d | 2026-08-20 | ## Resolved / archive @@ -505,3 +495,17 @@ Move resolved rows here with the resolution date and a one-line outcome. Keep th | #183 | task | Create Sentry metric alert for production DB span p95 > 500ms | Configured Sentry alert routing on critical worker spans | 2026-08-18 | | #FEWQZ5 | task | Therapy Compass has three shared-component stages left: Button call sites, card-recipes adoption, and page headers | Done — all three stages landed in PR #2122 (squash 092633eb). B3: 19 control-recipe call sites converted to the shared Button and commandControl/outlineControl/iconControl deleted. C: all 16 therapy card surfaces converged on card-recipes cardSurface, with the hero accent edge kept on the left as heroAccentEdge reading the shared --cat-accent. D: six page headers adopted the shared PageHeader; three h1 remain and are tracked separately as #D6G8TC. Verified by verify:cheap (673 files, 7276 tests), verify:ui (447 passed) and check:design-system-contract; ratchets fell (legacy shadow aliases 104 to 89, padding 53 to 52, gaps 30 to 28) and none rose. | 2026-08-18 | | #9DGA6R | task | Build packet B4: Docling worker shadow mode (WORKER_DOCUMENT_EXTRACTOR_MODE=legacy\|shadow) — authorised by the Gate B PASS of 2026-08-18 | Resolved 2026-08-19 by PR #2170 (branch claude/docling-worker-shadow-mode-b6fa17): typed WORKER_DOCUMENT_EXTRACTOR_MODE=legacy\|shadow (default legacy) + WORKER_SHADOW_EXTRACTION_COHORT_PERCENT (1-5, owner-approved 2) + WORKER_DOCLING_PYTHON_BIN; docling runs only after commitDocumentIndexGeneration on an index-quality-selected PDF cohort (tables / OCR / layout proxy), aggregate numbers-only record in documents.metadata.shadow_extraction via the existing metadata merge, no chunk/embedding/index/table-fact/document_index_quality writes, bounded 120 s / 40 pages / one process, fail-open (lost lease swallowed); docling venv + models provisioned in Dockerfile.worker from the Gate B lab lock (CI container build green); rollback WORKER_DOCUMENT_EXTRACTOR_MODE=legacy. Evidence: verify:pr-local heavy plan exit 0 (7292 tests), ingestion-worker-reviewer approve-with-nits (fixed). Both Gate B caveats carried in the PR body; enabling shadow in production is an operator Railway-variable step gated by docs/worker-deploy-runbook.md preconditions. | 2026-08-19 | +| #ZF006G | rec | Five independent SectionHeading implementations exist across modes with no shared recipe | Resolved 2026-08-21. Verified on main at 1cc0d2987: there is now exactly one SectionHeading definition, src/components/ui/section-heading.tsx, and the former independent implementations re-export it - dashboard-shell.tsx re-exports SectionHeading and SectionHeadingProps, formulation-ui.tsx and specifier-ui.tsx re-export SectionHeading. The fifth (therapy-compass/ui.tsx) was already deleted unreferenced. The stop rule was respected: the shared component lives in its own file, not folded into ui-primitives.tsx. Verification: grep for the SectionHeading declaration across src/ (one hit) plus the three re-export lines. | 2026-08-20 | +| #Q5JHBJ | task | Deploy the 20260818090000 schema_drift_snapshot v2 history probe (Phase 6.1) in an approved production window after Phase 4, triage the first unguarded no-statements report, and author fail-fast guard migrations for the pre-contract 2026-07-01..02 and 2026-07-12 rows | Closed 2026-08-20: every deliverable named in this row's own summary is complete and evidenced in its detail. The 20260818090000 schema_drift_snapshot v2 history probe was deployed in the owner-authorised production window, the first unguarded no-statements report was triaged (all fifteen versions classified), and the fail-fast guard migrations were authored and landed (six migrations, allowlist now 20 entries). Proof recorded in the row: full chain replay Applied 210/210 with CHAIN == MANIFEST, seven mutants raise and name their object, production and staging both green, and live-drift run 32251326536 reported No unexpected schema drift with all 20 history rows allowed. The row's detail already states the one residual is NOT this row's scope and is tracked separately as #M54C4N (the PGRST106 migration-history alignment step). Leaving the row in Open items made completed deployment and guard work re-dispatchable. Raised by the Codex review of PR #2206 (comment 3825543312). | 2026-08-20 | +| #DREDWA | rec | Ledger writer self-tests use only legacy numeric ids, which is why a Crockford-id lookup bug survived the ULID migration unnoticed | Resolved 2026-08-21. Verified on main at 1cc0d2987: Crockford/ULID-style id fixtures are now present in the shared ledger test surfaces - 12 matching occurrences in tests/outstanding-issues-writer.test.ts and 10 in scripts/outstanding-issues.mjs, including the all-digit case that defeats id-shape branching. The writer entry points are therefore driven through both id generations. Verification: grep count for the Crockford fixture markers in both files. | 2026-08-20 | +| #43SSS0 | rec | Three spring easing tokens in globals.css are dead: zero var() references and zero utility usage | Resolved 2026-08-21. Verified on main at 1cc0d2987: --spring-tight, --spring-bouncy and --spring-gentle have zero occurrences anywhere under src/, including src/app/globals.css, so the three dead theme tokens were deleted rather than wired. --ease-spring remains and is still referenced 8 times from globals.css, so the working precedent is intact. Verification: ripgrep over src/ for all three token names, no matches. | 2026-08-20 | +| #VZN8G3 | task | docling-lab-fixtures.v2: add unruled, merged-cell and rotated-header table fixtures to the Docling lab before any table-quality promotion argument | Closed as a duplicate 2026-08-20. #VZN8G3 and the earlier, still-open #BSBE9B are the same task: both require a docling-lab-fixtures.v2 corpus with unruled tables, merged/spanning cells and rotated headers, and both gate any docling table-quality promotion argument on that corpus existing. #BSBE9B is the canonical row and stays open; keeping both invites parallel implementations and conflicting closure state. The duplicate arose because inbox request a20fc4ce was queued without first checking the ledger for an existing row, and it was only visible once reconciliation allocated the id. Raised by the Codex review of PR #2206 (comment 3825543316). | 2026-08-20 | +| #E0N0QC | rec | scripts/probe-generation-quality.ts prints no answer text, so a blinded Gate E before/after read cannot come from probes | Resolved 2026-08-21. Verified on main at 1cc0d2987: scripts/probe-generation-quality.ts now exports formatAnswerSection(answerText, citations) and renders the trimmed answer text, falling back to '(empty answer)', so a blinded before/after Gate E read can be produced from probe output. Verification: grep for answerText in the probe script. Note the usual approval boundary still applies to running the probe against a provider. | 2026-08-20 | +| #EH9VA6 | rec | The serialized issues:reconcile operation has no interlock, and two reconcile PRs were open simultaneously on 2026-08-18 | Closed 2026-08-21 after verifying the cross-worktree interlock the original row asked for already exists. scripts/ledger-inbox.mjs implements two distinct protections, not one: reconcileLockPath()/RECONCILE_LOCK_NAME plus recoverReconcileTransaction() cover same-machine concurrency, and findUnmergedRemoteReconciliations() (line 428) covers the recorded 2026-08-18 failure of two reconcile PRs open simultaneously on GitHub. The remote check runs `git ls-remote --heads origin`, flags branches carrying unapplied records under docs/outstanding-issues-inbox/applied/, and names issue #EH9VA6 in both its docstring and its error text. assertSafeRemoteReconciliation() (line 603) is called at line 709 before any reconciliation work and refuses to reconcile on detection, so the gap the row described as unverified is in fact closed. Two residuals are deliberate and documented rather than outstanding: the remote check fails open when offline or when ls-remote errors, so offline workflows and tests are not blocked, and it can be overridden explicitly with --allow-concurrent/--force or ALLOW_CONCURRENT_RECONCILE=true, which downgrades the refusal to a warning. Neither is the silent same-machine-only blind spot the row was tracking. The earlier draft of this request queued an `update` saying only a filesystem lock existed; that was written from a partial read of the same file and would have made the canonical ledger advertise completed work as missing. | 2026-08-20 | +| #6K9YGQ | issue | Three standalone mode homes have no loading.tsx (/medications, /calculators, /dictionary) and the contract test cannot catch it | Resolved 2026-08-21. Verified on main at 1cc0d2987: loading.tsx now exists for all three standalone mode homes (medications, calculators and dictionary under src/app/(search-app)/), and tests/mode-home-loading-contract.test.ts now lists medications, calculators and dictionary in its route array, so the gap is visible to CI. Residual, deliberately not reopened: the contract test's route list is still hand-written rather than derived from standaloneModeHomePaths, which is a hygiene preference, not the reported invisibility. Verification: directory listing plus grep of the contract test. | 2026-08-20 | +| #0EKBGC | issue | Three mode homes override the canonical APP_MODE_ICON glyph (services, forms, dictionary) | Resolved 2026-08-21. Verified on main at 1cc0d2987: all three mode homes now derive their glyph from the canonical source - the services home imports appModeIcons and passes icon={appModeIcons.services}, the forms home passes icon={appModeIcons.forms}, and the dictionary home passes icon={appModeIcons.dictionary}. No local Users/FileText/BookOpen override remains on the home tile. Verification: grep for appModeIcons in each mode's home component. | 2026-08-20 | +| #TAQKCN | task | One hand-drawn SVG checkmark survived the Therapy Compass lucide sweep | Resolved 2026-08-21. Verified on main at 1cc0d2987: src/components/therapy-compass/screens/recommend-screen.tsx contains no | task | Three Therapy Compass h1 elements still sit outside PageHeader, and the patient-sheet builder renders two of them at once | Resolved 2026-08-21. Verified on main at 1cc0d2987: src/components/therapy-compass/ contains zero

) was not triggered. Verification: grep for ' | rec | check:drift clips table column diffs to 240 chars per side, so a wide-table column drift never names the column | Resolved 2026-08-21. Verified on main at 1cc0d2987: scripts/check-drift.ts now exports diffColumns, which builds name-keyed maps of the manifest and live column arrays and diffs per column rather than serialising and clipping each side to 240 characters. A late-alphabet column drift such as token_estimate on document_chunks can now be named. Verification: read of the diffColumns implementation in check-drift.ts. | 2026-08-20 | +| #5JK9FM | rec | PR template carries no RAG impact: guidance although pr-policy hard-blocks RAG-surface PRs without the line | Resolved 2026-08-21. Verified on main at 1cc0d2987: .github/pull_request_template.md now contains 4 occurrences of 'RAG impact', so the exact-format contract enforced by scripts/pr-policy.mjs ragImpactDeclared is visible where the body is authored. Verification: grep -c 'RAG impact' on the template. | 2026-08-20 | +| #HSSHRG | issue | The in-flight-CI guard closed as #145 does not cover merge-main syncs made outside sync:pr-branches — PR #2149 lost 8 of 9 CI cycles to self-inflicted cancellation | Resolved 2026-08-21. Verified on main at 1cc0d2987: scripts/guard-push.mjs now carries an explicit 'Guard 2: in-flight CI push guard (#HSSHRG)' block with inFlightCiVerdict() and findInFlightCiRuns(), and its header documents the cancel-in-progress waste this closes. Coverage therefore no longer depends on going through sync:pr-branches. Not verified from here: behaviour against a live PR with runs in flight, which needs GitHub access. Related row #TF6TPJ shares this root cause and should be re-checked against the same guard before it is closed separately. | 2026-08-20 |