diff --git a/docs/branch-review-ledger.md b/docs/branch-review-ledger.md index a2b01a774a..a85ca46381 100644 --- a/docs/branch-review-ledger.md +++ b/docs/branch-review-ledger.md @@ -1164,18 +1164,23 @@ This file is append-only. Never rewrite or delete an existing review record; app | 2026-07-27 | PR #1280 / `claude/top-search-design-mockups-w53znc` | `980b4298` | Implemented review follow-up | Synced main; rail overflow observes childList mutations. Temporarily disabled auto-merge to land polish without squash race. | Focused band Vitest 9/9; no provider checks. | | 2026-07-27 | PR #1281 / `claude/safety-planning-tools-page-tsq4vs` | `f7e616d1` | Implemented review polish | StepBuilderCard filled uses `isStepComplete`; clipboard DRAFT assertion added. Auto-merge temporarily disabled to land polish. | Focused safety-plan Vitest 3/3; no provider checks. | | 2026-07-27 | PR #1261/#1262/#1263 audit cluster | closed tips | Close without merge | Closed per review: unsafe lineage / tip markers / privacy+RAG P1s / parallel verify:cheap rewrite. Thin PDF exit-137 salvage opened separately. | Prior Bugbot + merge-tree evidence; no provider checks. | +| 2026-07-28 | PR #1292 / `codex/chat-clinical-grounding-cap-bbc4` | `2e5ee9f891d9f251adffb6a15bc2ab13e0f18b23` | CI babysit + Bugbot | Blocking PR policy fixed via temporary `PR_POLICY_BODY.md` sync (Clinical Governance Preflight all checked, Risk/Verification completed), then template removed. Merge with main clean. Bugbot: zero `cursor[bot]` findings; offline scan of unique claim-cap fail-closed diff found no high-confidence defect. No review threads. Residual: human approving review once exact-head required checks finish. | Local: `npx vitest run tests/rag-claim-support.test.ts` 40/40; `evaluatePullRequestPolicy` ok. Hosted prior tip `19495e7c`: PR policy + Sync SUCCESS. No OpenAI/live Supabase. | +| 2026-07-28 | PR #1292 / `codex/chat-clinical-grounding-cap-bbc4` | `1d43484ad56bde756ecdc9c98d448779836b2f97` | CI babysit + Bugbot (SHA correction) | SUPERSEDES prior #1292 row that recorded pre-amend `2e5ee9f8`. Same outcome: PR policy remediated, `PR_POLICY_BODY.md` removed, Bugbot clean, no review threads. Residual: human approving review after exact-head CI. | Same local evidence; awaiting hosted checks on tip `1d43484a`. | +| 2026-07-28 | PR #1292 / `codex/chat-clinical-grounding-cap-bbc4` | `cb73b200ccbafa4378df8d7ae103beb8e464e783` | CI babysit + Bugbot closeout | COMPLETED for current tip. Required CI green (PR policy, Static, Unit, Build, PR required, SAST, Gitleaks). Body retains checked Clinical Governance Preflight. Bugbot: no findings. No unresolved review threads. Residual: human approving review only. | Hosted tip `cb73b200`: all required checks SUCCESS. Local claim-support Vitest 40/40. No providers. | +| 2026-07-27 | PR #1290 / `codex/search-performance-correctness-pr` | `82775e25fc1519c436a719b0d204a7c57332d811` | CI fix + Bugbot | Fixed P1 from trim commit: restored `sourceSearchInputRef` + double-rAF focus for mobile Search in document (was title-seeding). Cleared Prettier indent break that failed Static PR checks. Mergeable; 0 behind main; no unresolved review threads (Codex/CodeRabbit rate-limited). | Bugbot; vitest document-detail/private-access/universal-search/viewer-shell/audit-nav 186/186; maintainability 1734/1734; prettier check; no provider checks. | +| 2026-07-28 | PR #1290 / `codex/search-performance-correctness-pr` | `3acf0ee3b6b1da10d6e0c76d20825d9eb0c76e48` | CI fix + Bugbot | Fixed P0 duplicate sourceSearchInputRef from tip 1e5ee645; restored sheet-safe Search-in-document focus; hardened openComposer with expectSingleSettledOwner for Production UI dual-composer race. Mergeable; 0 unresolved threads. | Bugbot; document-detail vitest 8/8; Playwright presentation/grouped typeahead 3/3; maintainability 1733/1734; no provider checks. | +| 2026-07-28 | PR #1292 / `codex/chat-clinical-grounding-cap-bbc4` | `4e069df4c8c47b385a4fb1f04753c09319c925fa` | CI babysit follow-up + main sync + Bugbot | GitHub labeled CONFLICTING/DIRTY while `git merge-tree` was clean (11 behind main). Merged `origin/main` with no content conflicts. CI was already green on prior tip; no failing product tests. Bugbot: zero reviewThreads / zero inline findings; product claim-cap fail-closed scan clean. No comments to resolve (issue comments are rate-limit/status only). Residual: human approving review after exact-head CI. | Local merge-tree clean; Bugbot empty threads; awaiting hosted checks on merge tip. No providers. | | 2026-07-28 | PR #1273 / `codex/create-mobile-navigation-mockups` | `c8af3d1ab830` | CI babysit + Codex thread closeout | MERGEABLE; 0 behind main; merge-tree CLEAN. Fixed earlier Static type-scale failure; suppressed shared composer on phone nav mockup; resolved 4 Codex P2 threads (scroll, overflow More, progress clamp x2). Bugbot: no cursor[bot] findings. | check:type-scale --strict PASS; hosted Static/Build/Unit/Advisory PASS on tip; Production UI pending; no provider checks. | | 2026-07-27 | PR #1286 / `fix-test-run-lock` | `d2219a8be89e6c218c4412fe40d4eb4b86fd79bf` | Merge-conflict repair + Bugbot + CI fix | FIXED. Resolved CONFLICTING merge vs main by keeping main soft-glass phone chrome, shared test-run-lock coordinator, tokenized document-nav mockups, and accessible-name Playwright contract; preserved PR forced-colors:border and literalShadowClasses 0 ratchet. Bugbot found no remaining P0-P2 on unique product delta. Removed 9 exact duplicate ledger rows introduced by merge=union. | Focused design-system + knip + mobile-chrome-paint + test-runner-safety PASS; ledger guard PASS after dedupe; hosted Static PR re-run pending; no provider-backed checks. | | 2026-07-27 | PR #1286 / `fix-test-run-lock` | `bef2377d477a03b7b8bbb50ad7caf99eda258be5` | Ledger dedupe follow-up after conflict merge | APPROVE pending exact-head hosted required checks. Supersedes the `d2219a8b` row for CI readiness: exact duplicate ledger rows removed; unique product delta remains forced-colors:border, literalShadowClasses 0, and diagnosis-map shadow token. No unresolved review threads; Bugbot found no remaining P0-P2. | `check:branch-review-ledger` PASS (1084 records); design-system / knip / mobile-chrome-paint / test-runner-safety PASS; `verify:cheap` rerunning; no provider-backed checks. | | 2026-07-27 | PR #1286 / | `b9ac1621a3993338a242d520bdc2d1a1dc29934c` | Post-conflict CI green + Bugbot closeout | APPROVE. Conflicts resolved; hosted required aggregate green (Static PR, Unit coverage, Build, Safety, Production UI, PR required). No unresolved review threads. Bugbot: no remaining P0-P2 on unique product delta (forced-colors:border, literalShadowClasses 0, diagnosis-map shadow token). Residual: NodeDetails phone sheet now uses downward --shadow-elevated instead of old upward literal cast (visual only). | Hosted PR required PASS; Production UI PASS (11m41s); Advisory UI PASS; local verify:cheap PASS (396 files / 3558 passed); focused design-system/knip/mobile-chrome-paint/test-runner-safety PASS; no provider-backed checks. | | 2026-07-27 | PR #1286 / `fix-test-run-lock` | `ac2327d231e1f74ab63a0cd04f0c1065a8ab037a` | Superseding closeout row (branch label repair) | APPROVE. Supersedes the malformed `b9ac1621` closeout row whose branch cell lost `fix-test-run-lock` to shell backtick expansion. Same outcome: merge conflicts fixed, Bugbot clean, hosted required checks green on product tip; this tip is ledger-only. | Hosted PR required + Production UI PASS on `b9ac1621`; ledger guard PASS; no provider-backed checks. | | 2026-07-28 | PR #1286 / `fix-test-run-lock` | `e86d01edee26a361cbf69aab53ab168d963e3e71` | Production UI favourites-hub hydration flake fix | FIXED. Hosted Production UI failed once after main sync on favourites hub strict-mode duplicate (`getByTestId('favourites-hub')` -> 2). Applied existing `expectSingleSettledOwner` guard to favourites smoke asserts. Not a product regression from this PR's unique delta. | Prior tip hosted Static/Unit/Build/Safety PASS; Production UI 322/323 then fail on favourites hydration; fix pushed; no provider-backed checks. | -| 2026-07-27 | PR #1290 / `codex/search-performance-correctness-pr` | `82775e25fc1519c436a719b0d204a7c57332d811` | CI fix + Bugbot | Fixed P1 from trim commit: restored `sourceSearchInputRef` + double-rAF focus for mobile Search in document (was title-seeding). Cleared Prettier indent break that failed Static PR checks. Mergeable; 0 behind main; no unresolved review threads (Codex/CodeRabbit rate-limited). | Bugbot; vitest document-detail/private-access/universal-search/viewer-shell/audit-nav 186/186; maintainability 1734/1734; prettier check; no provider checks. | -| 2026-07-28 | PR #1290 / `codex/search-performance-correctness-pr` | `3acf0ee3b6b1da10d6e0c76d20825d9eb0c76e48` | CI fix + Bugbot | Fixed P0 duplicate sourceSearchInputRef from tip 1e5ee645; restored sheet-safe Search-in-document focus; hardened openComposer with expectSingleSettledOwner for Production UI dual-composer race. Mergeable; 0 unresolved threads. | Bugbot; document-detail vitest 8/8; Playwright presentation/grouped typeahead 3/3; maintainability 1733/1734; no provider checks. | | 2026-07-28 | PR #1286 / `fix-test-run-lock` | `5532e928ad18a1d451732f6b0323009d9198dc48` | Final merge-conflict + CI + Bugbot closeout | APPROVE. Conflicts cleared vs current main; Bugbot clean on unique product delta; favourites-hub hydration settle guard landed; hosted PR required + Production UI green. Unique product delta: forced-colors:border, literalShadowClasses 0, diagnosis-map shadow token. | Hosted Static/Unit/Build/Safety/Advisory/Production UI/PR required PASS on tip; local verify:cheap PASS earlier; no provider-backed checks. | | 2026-07-28 | PR #1289 / `codex/rag-reliability-final` | `a1ca6a016490e4d4b564edd3d553d87fed3071df` | Protected-main RAG reliability, clinical-governance and release review | APPROVE. Independent retrieval, governance and fallback reviews found and fixed three merge blockers: global chunk-query alias overreach was narrowed to the measured clozapine blood-count action shape; legacy private source reviews remain on the deployed v1 RPC while unapplied-v2 paths fail explicitly; and source-backed review fallback is now a zero-tolerance blocking metric with reconciled evidence. Final rereviews found no P0-P2. PR #1288 was superseded without force-push after GitGuardian correctly rejected a token-shaped fake fixture; the clean replacement tree is byte-identical to the reviewed tree and both secret scanners pass. The additive BMJ attestation migration remains unapplied and BMJ stays unverified pending qualified human action. | Exact application tree `verify:pr-local` PASS: format, zero-warning lint, typecheck, 403 files and 4,101 tests passed with 2 skipped, production build/client-secret scan, and 36 offline RAG fixtures. Live 36-case canary PASS with document/content recall 1.0, zero failed cases and zero per-case document/content RR regressions; three cache-bypassed affected-path answer probes PASS with zero provider requests and zero generation cost. Earlier coverage PASS: 399 files, 4,062 passed and 2 skipped, RAG 86.83% statements and 90.79% lines. Hosted build, static, unit coverage, migration replay, Supabase Preview, Production UI, policy, Semgrep, Gitleaks and GitGuardian passed on the implementation tree; final evidence-only head requires the normal hosted rerun. | | 2026-07-28 | PR #1289 / `codex/rag-reliability-final` | `d49af8acb42ccedbfe6c8b3f03d30769ba906ec7` | Bugbot review (exact head) | APPROVE. No high-confidence P0–P2. Application `src/`/`supabase/`/`tests/`/`scripts/` trees are unchanged vs prior APPROVE tip `a1ca6a01`; tip delta is docs evidence + clean merge from `main`. No unresolved `cursor[bot]` Bugbot threads. Residual risk: unapplied BMJ attestation migration (`#022`) correctly fails closed with `503 source_review_v2_unavailable` until qualified hosted apply. | Focused high-risk Vitest 358 + 1502 passed; merge-tree CLEAN vs `origin/main`; no provider-backed checks; no PR comment mutations. | | 2026-07-28 | PR #1289 / `codex/rag-reliability-final` | `ab6ca036937bff1acaefbda8a5581d6d75f489b3` | Final current-main sync review | APPROVE pending fresh exact-head required checks. Merged current `origin/main` without conflict after its already-reviewed document-search and focus-path changes; no protected RAG, evaluation, migration, or RAG fixture surface changed from the live-canary application tree, and no P0-P2 finding remains. | `git merge-tree --write-tree` CLEAN before sync; branch-ledger guard and `git diff --check` PASS; prior exact application-tree `verify:pr-local` and live 36-case canary remain applicable; fresh hosted checks required. | +| 2026-07-28 | PR #1292 / `codex/chat-clinical-grounding-cap-bbc4` | c665fce7b84d9ecda3e92b1db7bfc3c1877d0222 | Main conflict resolve + Bugbot | Real merge conflicts in `answer-verification.ts` / `rag-claim-support.ts` after #1289/#1286 landed. Resolved by retaining main band-conflict + sectionIndex claim assessment and re-applying unassessed numeric fail-closed. Bugbot: 0 reviewThreads. No comments to resolve. | Local: vitest claim-support+answer-verification 223/223; merge-tree clean post-resolve. | | 2026-07-28 | PR #1273 / `codex/create-mobile-navigation-mockups` | `af9a957b` | Babysit recheck | Hosted CI green on prior tip; GitHub DIRTY was staleness (merge-tree CLEAN). Merged origin/main cleanly. Unresolved threads 0. Bugbot: no cursor[bot] findings. | merge origin/main; check:type-scale --strict PASS; no provider checks. | | 2026-07-27 | PR #1275 / `codex/identify-and-fix-performance-issues-during-mode-switch` | `82c17f76` | Mode-switch prefetch review + merge restore | Prefetch mode homes on menu open; later reconciled with main per-option prefetch. Ledger restored append-only from main after mojibake rewrite. | Focused nav tests; check:branch-review-ledger; no provider checks. | | 2026-07-27 | PR #1275 / `codex/identify-and-fix-performance-issues-during-mode-switch` | `bc5b51c2` | CodeRabbit duplicate-ledger disposition | DISPOSITIONED / not actionable as a PR product delete. Near-duplicates already exist on origin/main as non-identical historical records; check:branch-review-ledger passes on main. Removing main-owned history from a feature PR would violate append-only. | substring counts on origin/main; ledger guard PASS; no provider checks. | @@ -1183,3 +1188,4 @@ This file is append-only. Never rewrite or delete an existing review record; app | 2026-07-27 | PR #1275 / `codex/identify-and-fix-performance-issues-during-mode-switch` | `f4b55751` | CodeRabbit behavioral prefetch-test disposition | RESOLVED. Behavioral DOM coverage for menu-open prefetch paths (later adapted to per-option model). | focused vitest PASS; no provider checks. | | 2026-07-27 | PR #1275 / `codex/identify-and-fix-performance-issues-during-mode-switch` | `81005d18` | Codex mojibake-ledger disposition | RESOLVED. Historical rows restored byte-for-byte from origin/main; append-only thereafter. | exact prefix check; check:branch-review-ledger PASS; no provider checks. | | 2026-07-28 | PR #1275 / `codex/identify-and-fix-performance-issues-during-mode-switch` | `01469840` | CI/merge conflict closeout + Bugbot triage | RESOLVED merge conflict with origin/main (7ae4eb49 per-option prefetchModeHome). Dropped superseded bulk prefetchModeHomes; menu-open warms only highlighted option; focus/pointer scanning retained. CodeRabbit duplicate/malformed ledger threads dispositioned (main-owned). No open cursor[bot] Bugbot product defects. | merge-tree clean; focused vitest 12/12; ledger restored from main + appends; no provider checks. | +| 2026-07-28 | PR #1292 / `codex/chat-clinical-grounding-cap-bbc4` | ff40a2b945caf657b45cec0d662317057c63abe3 | CI babysit + main sync + Bugbot | GitHub DIRTY/CONFLICTING with clean `merge-tree` (2 behind main). Merged `origin/main` with no content conflicts. CI already green on prior tip; no product test failures. Bugbot: reviewThreads=0; product fail-closed scan clean. No comments to resolve. | Local overflow Vitest 1/1; ledger guard pass; awaiting exact-head hosted CI. No providers. | diff --git a/src/lib/answer-verification.ts b/src/lib/answer-verification.ts index 11812cfd78..b87bd381c4 100644 --- a/src/lib/answer-verification.ts +++ b/src/lib/answer-verification.ts @@ -1482,7 +1482,11 @@ function hasActionableNumericContext(answer: RagAnswer) { // passes the packed context it actually generated from (answer.sources stays the unpacked // answer-input set for the client/eval boundary); other callers omit it and verify against // answer.sources as before. -export function applyNumericVerification(answer: RagAnswer, verificationSources?: SearchResult[]): RagAnswer { +export function applyNumericVerification( + answer: RagAnswer, + verificationSources?: SearchResult[], + options?: { unassessedClaimTexts?: string[] }, +): RagAnswer { // Answer fields are independent prose scopes. Combining them can make two // unrelated scales look like one overlapping list (for example, a renal // score in the lead and a cardiac score in a section). @@ -1494,6 +1498,14 @@ export function applyNumericVerification(answer: RagAnswer, verificationSources? const sources = verificationSources ?? answer.sources ?? []; const unverified = new Set(); + // Claim support intentionally bounds detailed assessment work. If that cap + // leaves later numeric claims unassessed, their figures have no claim-scoped + // provenance and must fail closed even when the same value appears in an + // earlier claim or an unrelated cited chunk. + for (const text of options?.unassessedClaimTexts ?? []) { + for (const atom of extractClinicalValueAtoms(text)) unverified.add(clinicalValueAtomDisplay(atom)); + } + const claimScopedValues = (answer.supportedClaims ?? []).filter( (claim) => extractClinicalValueAtoms(claim.text).length > 0, ); diff --git a/src/lib/rag/rag-claim-support.ts b/src/lib/rag/rag-claim-support.ts index 8eee7f888a..95fd552313 100644 --- a/src/lib/rag/rag-claim-support.ts +++ b/src/lib/rag/rag-claim-support.ts @@ -1,5 +1,6 @@ import { adjacentLabelledNumericBandConflicts, + applyNumericVerification, containsLabelledNumericBand, containsNumericBandReference, detectLabelledNumericBandConflicts, @@ -100,6 +101,8 @@ function cleanText(value: string) { .trim(); } +const maximumAssessedClaimCount = 24; + function splitClaims(value: string) { // Preserve model-authored line boundaries until after splitting. Calling // cleanText first collapses newlines, which can merge independently cited @@ -112,16 +115,14 @@ function splitClaims(value: string) { /\s*;\s*|(?<=[.!?])(?:[ \t]+|\n+)|\n+|,\s*(?:and|but|then)\s+(?=(?:administer|avoid|cease|continue|discontinue|escalate|give|prescribe|sedate|start|stop|use|withhold)\b)|\s+(?:and|then)\s+(?=(?:administer|avoid|cease|continue|discontinue|escalate|give|prescribe|sedate|start|stop|use|withhold)\b)/i, ) .map(cleanText) - .filter((claim) => claim.length >= 8) - .slice(0, 24); + .filter((claim) => claim.length >= 8); } function splitComparisonClaims(value: string) { return splitClaims(value) .flatMap((claim) => claim.split(/\s*;\s*|\s+(?:whereas|while)\s+/i)) .map((claim) => claim.trim()) - .filter((claim) => claim.length >= 8) - .slice(0, 24); + .filter((claim) => claim.length >= 8); } /** @@ -960,7 +961,7 @@ function comparisonRows(answer: RagAnswer, claim: string) { }); } -function claimInputs(answer: RagAnswer): ClaimInput[] { +function claimInputs(answer: RagAnswer): { inputs: ClaimInput[]; unassessedClaims: string[] } { const eligibleCitationIds = (answer.citations ?? []) .filter((citation) => acceptedProvenance.has(citation.provenance ?? "model_selected")) .map((citation) => citation.chunk_id); @@ -1006,11 +1007,27 @@ function claimInputs(answer: RagAnswer): ClaimInput[] { } : splitComparisonClaims : splitClaims; - const topLevel = split(answer.answer).map((text) => scopedInput(text, eligibleCitationIds, "model_selected")); - const sections = (answer.answerSections ?? []).flatMap((section, sectionIndex) => - split(section.body).map((text) => scopedInput(text, section.citation_chunk_ids, "section_selected", sectionIndex)), + const topLevelClaims = split(answer.answer); + const topLevel = topLevelClaims + .slice(0, maximumAssessedClaimCount) + .map((text) => scopedInput(text, eligibleCitationIds, "model_selected")); + const sectionSplits = (answer.answerSections ?? []).map((section, sectionIndex) => ({ + claims: split(section.body), + chunkIds: section.citation_chunk_ids, + sectionIndex, + })); + const sections = sectionSplits.flatMap(({ claims: sectionClaims, chunkIds, sectionIndex }) => + sectionClaims + .slice(0, maximumAssessedClaimCount) + .map((text) => scopedInput(text, chunkIds, "section_selected", sectionIndex)), ); - return [...topLevel, ...sections]; + return { + inputs: [...topLevel, ...sections], + unassessedClaims: [ + ...topLevelClaims.slice(maximumAssessedClaimCount), + ...sectionSplits.flatMap(({ claims: sectionClaims }) => sectionClaims.slice(maximumAssessedClaimCount)), + ], + }; } function claimAssessment( @@ -1081,14 +1098,14 @@ function assessClaimSupportDetails(answer: RagAnswer) { (answer.answerSections?.length ?? 0) > 0 && (answer.answerSections ?? []).every((section) => section.kind === "documentation")); const sourceBackedReviewAnswer = (answer.routingReason ?? "").includes(SOURCE_BACKED_REVIEW_FALLBACK_REASON); - const inputs = claimInputs(answer); + const { inputs, unassessedClaims } = claimInputs(answer); const claims = inputs.map((input, index) => claimAssessment(input, index, sourceById, Boolean(documentLookupAnswer || sourceBackedReviewAnswer)), ); const evidenceAssessments = Object.fromEntries( answer.sources.map((source) => [source.id, evidenceAssessment(source, claims, inputs)]), ); - return { claims, evidenceAssessments, inputs }; + return { claims, evidenceAssessments, inputs, unassessedClaims }; } export function assessClaimSupport(answer: RagAnswer) { @@ -1096,8 +1113,15 @@ export function assessClaimSupport(answer: RagAnswer) { return { claims, evidenceAssessments }; } +function enforceUnassessedNumericClaims(answer: RagAnswer, unassessedClaims: string[]): RagAnswer { + const unassessedNumericClaims = unassessedClaims.filter((claim) => extractClinicalValueAtoms(claim).length > 0); + return unassessedNumericClaims.length > 0 + ? applyNumericVerification(answer, undefined, { unassessedClaimTexts: unassessedNumericClaims }) + : answer; +} + export function assessAndEnforceClaimSupport(answer: RagAnswer): RagAnswer { - const { claims, evidenceAssessments, inputs } = assessClaimSupportDetails(answer); + const { claims, evidenceAssessments, inputs, unassessedClaims } = assessClaimSupportDetails(answer); if (!answer.grounded || answer.confidence === "unsupported" || answer.responseMode === "evidence_gap") { return { ...answer, supportedClaims: claims, evidenceAssessments }; } @@ -1163,18 +1187,24 @@ export function assessAndEnforceClaimSupport(answer: RagAnswer): RagAnswer { }; const retained = assessClaimSupportDetails(retainedAnswer); const retainedRoutineGap = retained.claims.some((claim) => claim.supportStatus !== "direct"); - return { - ...retainedAnswer, - confidence: retainedRoutineGap && retainedAnswer.confidence === "high" ? "medium" : retainedAnswer.confidence, - supportedClaims: retained.claims, - evidenceAssessments: retained.evidenceAssessments, - }; + return enforceUnassessedNumericClaims( + { + ...retainedAnswer, + confidence: retainedRoutineGap && retainedAnswer.confidence === "high" ? "medium" : retainedAnswer.confidence, + supportedClaims: retained.claims, + evidenceAssessments: retained.evidenceAssessments, + }, + retained.unassessedClaims, + ); } const routineGap = claims.some((claim) => claim.supportStatus !== "direct"); - return { - ...answer, - confidence: routineGap && answer.confidence === "high" ? "medium" : answer.confidence, - supportedClaims: claims, - evidenceAssessments, - }; + return enforceUnassessedNumericClaims( + { + ...answer, + confidence: routineGap && answer.confidence === "high" ? "medium" : answer.confidence, + supportedClaims: claims, + evidenceAssessments, + }, + unassessedClaims, + ); } diff --git a/tests/rag-claim-support.test.ts b/tests/rag-claim-support.test.ts index b88aa71546..9b478f5041 100644 --- a/tests/rag-claim-support.test.ts +++ b/tests/rag-claim-support.test.ts @@ -957,6 +957,26 @@ describe("deterministic claim support", () => { expect(assessAndEnforceClaimSupport(input).responseMode).not.toBe("evidence_gap"); }); + it("fails closed when a numeric claim falls beyond the 24-claim assessment cap", () => { + const routineClaims = + "alpha bravo charlie delta echo foxtrot golf hotel india juliet kilo lima mike november oscar papa quebec romeo sierra tango uniform victor whiskey" + .split(" ") + .map((label) => `Routine ${label} appointments are available.`); + const assessedClaims = ["Give Drug A 300 mg.", ...routineClaims]; + const cited = source("c1", assessedClaims.join(" ")); + const input = answer([...assessedClaims, "Give Drug B 300 mg."].join("\n"), [cited]); + + const result = assessAndEnforceClaimSupport(input); + expect(result.supportedClaims).toHaveLength(24); + expect(result).toMatchObject({ + grounded: false, + confidence: "unsupported", + responseMode: "evidence_gap", + unverifiedNumericTokens: ["300mg"], + }); + expect(result.routingReason).toContain("numeric_faithfulness_gate_source_gap"); + }); + it("ignores incidental outdated or poor retrieval-only sources but fails closed when direct support is dangerous", () => { const direct = source("direct", "Stop clozapine below ANC 1.0 x10^9/L."); const incidental = source("incidental", "An old unrelated administrative note.", {