Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions docs/branch-review-ledger.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -329,8 +329,8 @@ Records before 2026-07-28 were written by hand and had drifted: 146 lines carrie
| 2026-07-31 | claude/root-dir-coverage-gate-v2 | 398660144d93aeefc2e5649c156948a68925cb64 | docs:check-index repo-root coverage, stale script counts, ledger correction | MERGED as PR #1458 (squash 907fd9f4a). Root-directory coverage pass for docs:check-index, red-then-green proven (flagged .cursor/.design-sync/.vscode, then 49 entries vs 31). Main landed an equivalent pass independently in #1480, so the two overlapped; no duplication reached main. Row not recorded at the time - appended retrospectively | verify:cheap exit 0, 435 test files / 4574 tests pass; codebase-index-coverage 10/10 incl 4 new root cases; eslint clean; docs gates green; prettier clean |
| 2026-07-31 | claude/pre-commit-fail-open | 7b96a09b8500adc917cf5549b1c61142b2244b39 | pre-commit hook fail-open when the inventory script is absent | MERGED as PR #1494 (squash 387c3b653). Resolves ledger #153: core.hooksPath is absolute to the primary checkout, so the hook ran in worktrees lacking scripts/update-docs-inventory.mjs and aborted with MODULE_NOT_FOUND. Guard drops the inventory task and re-checks the all-tasks-empty exit; grep carries \|\| true because set -e treats a fully-filtering grep as failure | isolated-repo probe with the script genuinely absent: prints skipping inventory sync, commit succeeds; sh -n clean; no-op when the script is present; prettier does not parse shell so format:check skips it |
| 2026-07-31 | claude/ledger-relanding | 30ec06964e4235d9f0b4bb782f357e6b4fb59430 | re-land the three session findings lost when PR #1490 was closed | MERGED as PR #1508 (squash 7b551abc4). Ledger-only: #151 corrects the claim that CI is unreadable (PAT has Actions:read though not Checks:read), #152 re-lands the at-risk worktree inventory with the four preservation snapshots, #153 archives the hook fix. Verified landed by content on main, not by PR state or row id | CI, PR Policy, PR mergeability, SAST, Secret Scan all completed/success via the Actions API; check:outstanding-issues 151 rows 45 open unique ids next-id=154; docs:check-links 1414 refs; prettier clean |
| 2026-07-31 | claude/warning-consolidation-mockups-09jyj7 | 7b41fcf581085872da76270b109e2795c6940677 | PR #1437 warning consolidation mockups reopen prep | ready-closed: main merged clean; follow-ups renumbered #155-#157; bugbot P2s fixed; origin insteadOf false-positive fixed; verify:pr-local green (444/4646) | verify:pr-local;check:outstanding-issues;merge-tree:clean;pr-bugbot;diff-review |
| 2026-07-31 | claude/warning-consolidation-mockups-09jyj7 | b02cfc9258446f6f46bb6acfadd4e978950865c2 | PR #1437 warning consolidation mockups reopen prep | ready-closed at tip (ledger row + prior fixes); PR remains CLOSED; body update attempted | verify:pr-local@7b41fcf5;merge-tree:clean |
| 2026-07-31 | codex/complete-repository-maturity-programme | cb07a6c3698bc5683a0783555032490cbf09b674 | PR #1472 reopen prep | approved-with-notes | verify:pr-local green (444 files / 4645 tests); maintainability 4336/4351; focused RAG 730; bugbot none; CodeRabbit baseline fixed; main merged clean; PR left CLOSED |
| 2026-07-31 | codex/complete-repository-maturity-programme | 84fdfd72a5d23e79798be85ffee2dda4f6f6e94a | PR #1472 reopen prep | approved-with-notes | supersede cb07a6c3: tip is ledger-only after approved reopen prep; branch ready; PR remains CLOSED (GitHub freezes closed PR head until reopen) |
| 2026-07-30 | PR-1471 | ec6a8683c923cebbf3c028d6e2a89ee0efe32f18 | PR #1471 Therapy Compass payload | PASS: compact browse index preserves card fields and first-sentence best-used-for data while the search route explicitly loads the full catalogue | therapy index current at 205 records; focused Vitest 3 files, 20 tests passed; typecheck passed; provider-backed Lighthouse not run |
| 2026-07-30 | pr/1471 | fc6b890ad987ef107701d4ad05e4adf7556759d9 | Therapy browse payload and full-prose search preservation | no findings | 13 focused tests; therapy index guard; ledger guard |
| 2026-07-31 | codex/complete-and-merge-p2-tasks-to-main | bec88721b04054eda59372cf3e0c14c771e25e3b | PR #1471 Therapy Compass browse payload | PASS: origin/main synced (clean merge-tree; GitHub DIRTY was staleness); no Bugbot/CodeRabbit actionable threads; no P0-P2 findings; pathways compact-index fetch guard added; PR left CLOSED for reopen | node scripts/build-therapies-index.mjs --check (205); vitest therapy files 14 passed; typecheck passed; merge origin/main clean; no unresolved review threads |
Expand DownExpand Up@@ -536,3 +536,5 @@ Records before 2026-07-28 were written by hand and had drifted: 146 lines carrie
| 2026-07-31 | codex/chat-frontend-skill-selection-0978 | 7cf0505be4423f6856e45a6a87e9433fcae72462 | PR #1460 review+bugbot+fix | no findings; sync cleared GitHub DIRTY (merge-tree was behind-but-clean); tip delta ledger-only; product brace-expansion already on main via #1456; no Bugbot/actionable threads | merge-tree clean; check:branch-review-ledger PASS; required CI pending after sync push; prior missing checks while DIRTY not green |
| 2026-07-31 | codex/chat-frontend-skill-selection-0978 | c436387c49d37e65bca49429d2b74fffe61c0bfd | PR #1460 review+bugbot+fix | no findings; second sync after main advanced mid-pass; merge-tree clean; tip delta ledger-only; product already on main via #1456; no Bugbot/actionable threads | merge-tree clean; check:branch-review-ledger PASS; required CI expected after sync push |
| 2026-07-31 | codex/complete-and-merge-p2-tasks-to-main | 1761a100464992486fc7979a448c7a9773bd3e55 | PR #1471 review+bugbot+fix+heavy | PASS: synced origin/main (behind-but-clean); deep review+Bugbot no P0-P2; 0 threads; therapy index contracts OK; verify:cheap + verify:pr-local green; PR body RAG impact accurate | merge-tree clean; 0 behind; therapy --check 205; vitest therapy 14 passed; verify:cheap 445 files/4661 tests; verify:pr-local runtime+lint+typecheck+test+build+rag fixtures; prior tip ea92b37 CI pr-required green |
| 2026-07-31 | claude/warning-consolidation-mockups-09jyj7 | 7b41fcf581085872da76270b109e2795c6940677 | PR #1437 warning consolidation mockups reopen prep | ready-closed: main merged clean; follow-ups renumbered #155-#157; bugbot P2s fixed; origin insteadOf false-positive fixed; verify:pr-local green (444/4646) | verify:pr-local;check:outstanding-issues;merge-tree:clean;pr-bugbot;diff-review |
| 2026-07-31 | claude/warning-consolidation-mockups-09jyj7 | b02cfc9258446f6f46bb6acfadd4e978950865c2 | PR #1437 warning consolidation mockups reopen prep | ready-closed at tip (ledger row + prior fixes); PR remains CLOSED; body update attempted | verify:pr-local@7b41fcf5;merge-tree:clean |
2 changes: 1 addition & 1 deletion docs/codebase-index.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -123,7 +123,7 @@ domain-extracted directory; imported as `@/lib/rag/rag*`). Other modules below r
| `rag.ts` | Main answer pipeline orchestrator |
| `rag-routing.ts`, `rag-provider.ts`, `rag-answer-text.ts`, `smart-rag-api.ts` | Model routing, provider modes, API surface |
| `rag-contracts.ts`, `rag-answer-support.ts`, `rag-query-guard.ts` | Shared RAG contracts and pure answer/query policy |
| `rag-evidence-gates.ts`, `rag-coverage-gate.ts`| Evidence-sufficiency predicates and the fast-path evidence coverage gate |
| `rag-evidence-gates.ts`, `rag-coverage-gate.ts`, `rag-second-stage.ts` | Evidence predicates, fast-path coverage gating, and second-stage ranking |
| `rag-hydration.ts` | Per-request hydration: document ranking metadata, cached index quality, page visual evidence |
| `rag-cache.ts`, `rag-retrieval-variants.ts` | Bounded caches and retrieval variants |
| `clinical-search.ts`, `clinical-query-mode.ts`, `retrieval-selection.ts` | Query modes and retrieval selection |
Expand Down
6 changes: 6 additions & 0 deletions docs/maturity-backlog-workorders.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -92,6 +92,12 @@ structural change, not a single mixed PR.
`rag.ts`: it is pipeline orchestration that calls the metadata/visual hydration and
second-stage rerank helpers, so moving it would need a runtime back-edge to `rag.ts`. The
hydration cluster is the separate later extraction (`rag-hydration.ts`).
- **Progress (#086, X3 second-stage extraction, PR #1472):** extracted the cohesive
second-stage reranking unit into `src/lib/rag/rag-second-stage.ts`: engagement policy,
score adjustment, document diversity, subject-match protection, and its retrieval-layer
telemetry moved together without changing the public `@/lib/rag/rag` export. `rag.ts`
remains the retrieval orchestrator and calls the extracted unit at the same pipeline
points (4,543 → 4,351; budget ratcheted to 4,351).
- **Progress (`DocumentViewer.tsx`):** extracted the cohesive leaf modules into
`src/components/document-viewer/` — shared row `types.ts`, `source-panels.tsx` (summary
profile, high-yield summary, source images/tables, pinned evidence, indexed-text panel), the
Expand Down
7 changes: 3 additions & 4 deletions scripts/check-maintainability-budgets.mjs
Original file line numberDiff line numberDiff line change
Expand Up@@ -5,10 +5,9 @@ const budgets = new Map([
// Chrome ownership/reporting lives in use-dashboard-chrome-coordinator; keep
// the reclaimed monolith budget so it cannot silently drift back to 4160.
["src/components/ClinicalDashboard.tsx", 4140],
// The evidence coverage gate lives in rag-coverage-gate and per-request
// hydration in rag-hydration; keep the reclaimed budget so it cannot silently
// drift back to 5030.
["src/lib/rag/rag.ts", 4543],
// Evidence coverage, per-request hydration, and second-stage ranking live in
// focused rag modules; keep the reclaimed budget so it cannot silently drift back.
["src/lib/rag/rag.ts", 4351],
["src/components/DocumentViewer.tsx", 1734],
["supabase/functions/indexing-v3-agent/index.ts", 2191],
]);
Expand Down
199 changes: 199 additions & 0 deletions src/lib/rag/rag-second-stage.ts
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,199 @@
import { rankingConfig } from "@/lib/ranking-config";
import { visualEvidenceUnitTypes } from "@/lib/rag/rag-evidence-gates";
import type { SearchTelemetry } from "@/lib/rag/rag-contracts";
import type { RagQueryClass, SearchResult } from "@/lib/types";

// Extracted from rag.ts (maturity X3): second-stage engagement, scoring, and
// telemetry. The implementation is unchanged; rag.ts remains the orchestrator.

const tableVisualEvidenceUnitTypes = new Set([
"table_fact",
"table_threshold",
"medication_chart_row",
"risk_matrix_cell",
]);

/** Layer top score. */
export function layerTopScore(results: SearchResult[]) {
return Number(Math.max(0, ...results.map((result) => result.hybrid_score ?? result.similarity ?? 0)).toFixed(4));
}

/** Record retrieval layer. */
export function recordRetrievalLayer(
telemetry: SearchTelemetry,
layer: string,
count: number,
options: { latencyMs?: number; topScore?: number } = {},
) {
telemetry.retrieval_layer_counts = {
...(telemetry.retrieval_layer_counts ?? {}),
[layer]: count,
};
if (typeof options.latencyMs === "number") {
telemetry.retrieval_layer_latencies_ms = {
...(telemetry.retrieval_layer_latencies_ms ?? {}),
[layer]: Math.max(0, Math.round(options.latencyMs)),
};
}
if (typeof options.topScore === "number") {
telemetry.retrieval_layer_top_scores = {
...(telemetry.retrieval_layer_top_scores ?? {}),
[layer]: Number(Math.max(0, options.topScore).toFixed(4)),
};
}
}

/** Should use second stage rerank. */
function shouldUseSecondStageRerank(queryClass: RagQueryClass | undefined, results: SearchResult[], topK: number) {
if (results.length <= 1) return false;
const topScore = Math.max(0, results[0]?.hybrid_score ?? results[0]?.similarity ?? 0);
const secondScore = Math.max(0, results[1]?.hybrid_score ?? results[1]?.similarity ?? 0);
const topScoresClose = Math.abs(topScore - secondScore) <= 0.04;
const hasVisualEvidence = results.some((result) => visualEvidenceUnitTypes.has(result.index_unit?.unit_type ?? ""));
const hasTableVisualEvidence = results.some((result) =>
tableVisualEvidenceUnitTypes.has(result.index_unit?.unit_type ?? ""),
);
if (queryClass === "table_threshold" || queryClass === "medication_dose_risk") {
return hasVisualEvidence || hasTableVisualEvidence || topScoresClose;
}
if (queryClass === "comparison") return results.length > topK || topScoresClose;
return topScoresClose && hasVisualEvidence;
}

/** Second stage score. */
function secondStageScore(result: SearchResult, queryClass: RagQueryClass | undefined, index: number) {
const baseRankScore =
result.score_explanation?.rankScore ??
result.score_explanation?.preClampFinalScore ??
result.score_explanation?.finalScore ??
result.hybrid_score ??
result.similarity ??
0;
let adjustment = 0;
const unitType = result.index_unit?.unit_type ?? "";
const source = result.index_unit?.metadata?.source;
const sourceQuality = Number(result.index_unit?.quality_score ?? 0.65);
const doseAmountText = `${result.section_heading ?? ""} ${result.content} ${(result.images ?? [])
.map((image) => `${image.caption ?? ""} ${image.tableTextSnippet ?? ""} ${image.tableTitle ?? ""}`)
.join(" ")} ${(result.table_facts ?? [])
.map(
(fact) => `${fact.table_title ?? ""} ${fact.row_label ?? ""} ${fact.threshold_value ?? ""} ${fact.action ?? ""}`,
)
.join(" ")}`;
const hasDoseAmount = /\b\d+(?:\.\d+)?\s?(?:mg|mcg|microgram|micrograms)\b/i.test(doseAmountText);
const w = rankingConfig.secondStage;
adjustment += Math.max(0, w.positionBase - index * w.positionStep);
if (result.memory_cards?.length && (queryClass === "broad_summary" || queryClass === "comparison"))
adjustment += w.memorySummaryBoost;
if (queryClass === "document_lookup" && (result.match_explanation?.titleHit || result.match_explanation?.labelHit))
adjustment += w.documentLookupTitleBoost;
if ((queryClass === "table_threshold" || queryClass === "medication_dose_risk") && result.table_facts?.length)
adjustment += w.tableThresholdEvidenceBoost;
if (queryClass === "medication_dose_risk" && hasDoseAmount) adjustment += w.doseAmountBoost;
if (tableVisualEvidenceUnitTypes.has(unitType)) adjustment += w.tableVisualBoost;
else if (visualEvidenceUnitTypes.has(unitType)) adjustment += w.visualBoost;
if (source === "visual_intelligence")
adjustment += Math.min(
w.visualIntelligenceMax,
Math.max(0, sourceQuality - w.visualIntelligencePivot) * w.visualIntelligenceSlope,
);
if (result.source_metadata?.document_status === "outdated") adjustment -= w.outdatedPenalty;
// D4: ships 0 (no-op) — activate via RAG_RANKING_CONFIG only behind a green golden eval.
if (result.source_metadata?.document_status === "unknown") adjustment -= w.unknownCurrentnessPenalty;
if (result.source_metadata?.extraction_quality === "poor") adjustment -= w.poorExtractionPenalty;
if (
result.indexing_quality?.quality_score !== undefined &&
result.indexing_quality.quality_score < w.lowIndexQualityThreshold
)
adjustment -= w.lowIndexQualityPenalty;
return { rankScore: baseRankScore + adjustment, adjustment };
}

/** Apply second stage rerank if needed. */
export function applySecondStageRerankIfNeeded(args: {
queryClass?: RagQueryClass;
results: SearchResult[];
telemetry: SearchTelemetry;
topK: number;
}) {
if (!shouldUseSecondStageRerank(args.queryClass, args.results, args.topK)) return args.results;
const startedAt = Date.now();
// CI-16 document diversity: subtract a demotion from each EXTRA chunk of a document that
// has already appeared higher up, so a single doc's sibling chunks can't crowd out other
// documents. Applied AFTER the additive-boost floor so it can actually lower the effective
// rank. Keep this separate from the selection-rescue floor in retrieval-selection.ts.
const seenPerDocument = new Map<string, number>();
const reranked = args.results
.map((result, index) => {
const secondStage = secondStageScore(result, args.queryClass, index);
let rankScore = secondStage.rankScore;
let confidenceAdjustment = secondStage.adjustment;
const releasedHybridScore = result.hybrid_score ?? result.similarity ?? 0;
let releaseRankScore = Math.max(
releasedHybridScore,
(result.score_explanation?.finalScore ?? result.hybrid_score ?? result.similarity ?? 0) +
secondStage.adjustment,
);
const priorOccurrences = seenPerDocument.get(result.document_id) ?? 0;
seenPerDocument.set(result.document_id, priorOccurrences + 1);
if (rankingConfig.documentDiversityPenalty > 0 && priorOccurrences > 0) {
const diversityPenalty = Math.min(
rankingConfig.documentDiversityPenaltyCap,
rankingConfig.documentDiversityPenalty * priorOccurrences,
);
rankScore -= diversityPenalty;
confidenceAdjustment -= diversityPenalty;
releaseRankScore -= diversityPenalty;
}
const selectionReasons = result.match_explanation?.reasons ?? [];
const clinicalSubjectRequired = selectionReasons.includes("retrieval_required_signal:clinical_subject");
const clinicalSubjectMatched = selectionReasons.includes("retrieval_signal:clinical_subject");
if (clinicalSubjectRequired && !clinicalSubjectMatched) {
// A wrong-medication chunk can carry attractive numeric dose/monitoring signals. Keep it
// available at its released hybrid strength, but do not let second-stage evidence boosts
// promote it above chunks that contain the medication subject requested by the query.
releaseRankScore = Math.min(releaseRankScore, releasedHybridScore);
}
const finalScore = Math.min(
1,
Math.max(
0,
(result.score_explanation?.finalScore ?? result.hybrid_score ?? result.similarity ?? 0) +
confidenceAdjustment,
),
);
return {
rankScore,
result: {
...result,
score_explanation: result.score_explanation
? {
...result.score_explanation,
rankScore: Number(rankScore.toFixed(4)),
releaseRankScore: Number(releaseRankScore.toFixed(4)),
preClampFinalScore: Number(rankScore.toFixed(4)),
finalScore: Number(finalScore.toFixed(4)),
}
: result.score_explanation,
match_explanation: {
...result.match_explanation,
reasons: Array.from(new Set([...(result.match_explanation?.reasons ?? []), "second_stage_rerank"])),
},
},
};
})
.sort((left, right) => right.rankScore - left.rankScore || left.result.id.localeCompare(right.result.id))
.map(({ result }, index) =>
result.score_explanation
? { ...result, score_explanation: { ...result.score_explanation, finalRank: index + 1 } }
: result,
);
args.telemetry.second_stage_rerank_used = true;
args.telemetry.second_stage_rerank_latency_ms =
(args.telemetry.second_stage_rerank_latency_ms ?? 0) + Date.now() - startedAt;
recordRetrievalLayer(args.telemetry, "second_stage_rerank", reranked.length, {
latencyMs: Date.now() - startedAt,
topScore: layerTopScore(reranked),
});
return reranked;
}
Loading
Loading