Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | f4bc9033d5482922de60681d764eb5e6c6ebb73f | rag-candidate-sources.ts signal-row Zod contracts (ledger #212 tranche 2) + 2 ledger captures | Approved — 3 unchecked RPC-result casts replaced with validated assertions; similarity_origin item deliberately excluded as a clinical-output behaviour change; no ranking/ordering/scoring logic touched | verify:pr-local all 18 steps green, zero failures; contract + rag-imputation-contract pins 21 passed; typecheck exit 0 |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 690204f669db3be9995b6c658ad2eb35befbdace | PR #1981 base sync | Merged main 6f7b7deefaf7e0cd062b748f18fc6ca8988093f6 into the reviewed PR head; merge tree was clean. | git merge-tree --write-tree exact-head main: clean; git diff --check; ledger guards. |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 87a8886f1f761bdab92324e0d5ea5e11d3111bb9 | review-and-fix | P1 CI blocker fixed: formatted retrieval row contract test; no additional P0-P2 findings in adversarial review; merged latest main | targeted Vitest 134 pass; RAG fixtures 36 pass; offline RAG 579 pass; issue and ledger guards pass; Prettier pass |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 18253c2bc555d424e83f987712cf528dd3910e3f | Carry validated optional RAG signal degradation fix through main d301d8f4 | fixed | git diff --check; static RAG contract inspection; ledger and issue guards; tests blocked without node_modules |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 671b0b99f7fdd33e83e5fa55a29470690c9243f2 | RAG row-contract tranche 2 P2: unconstrained JSON provenance acceptance | Fixed P2 — index-unit source_span and metadata accept all JSON allowed by the database; non-object provenance is safely omitted from record-only downstream consumers | manual adversarial review; focused scalar/array contract regression added; git diff --check; ci-change-scope self-test; ledger/inbox/outstanding/discipline guards passed; targeted Vitest blocked: node_modules/vitest absent |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 1a59ce0b128fbdabb9a24c3ea95c1123299693f3 | RAG row-contract CI follow-up: optional-signal assertion narrowing | Fixed P1 — explicit local mismatch catches preserve both logged optional-signal degradation and TypeScript row narrowing | CI typecheck failure reproduced from exact-head log; git diff --check; ci-change-scope self-test; ledger/inbox/outstanding/discipline guards passed; local Vitest unavailable because node_modules/vitest is absent |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 690204f669db3be9995b6c658ad2eb35befbdace | PR #1981 retrieval row contract formatter follow-up | Collapsed a formatter-stable candidate-source import after the exact-head changed-file format gate failed; retained the validated row-shape assertions. | git diff --check; ledger/outstanding/branch-ledger/ledger-discipline guards; ci-change-scope self-test passed; npm test -- tests/rag-retrieval-row-contract.test.ts unavailable: node_modules/vitest/vitest.mjs absent. |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 20304ddc703fdc6913e1b62ad532b552ff919dac | RAG row-contract tranche 2: optional signal shape-mismatch degradation; merged main | Approved with P1 fix — malformed optional embedding-field/index-unit rows are logged and degrade to no signal candidates, preserving chunk retrieval; no ranking, ordering, or clinical-output contract changed | manual adversarial control-flow review; git diff --check; ci-change-scope self-test; ledger/inbox/outstanding/discipline guards passed; targeted Vitest blocked: node_modules/vitest absent |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
| 2026-08-15 | claude/rag-zod-hardening-tranche2 | 690204f669db3be9995b6c658ad2eb35befbdace | RAG signal-row formatter follow-up | Formatted the signal-row regression assertion reported by changed-file formatting. Targeted Vitest unavailable because this isolated worktree has no node_modules/vitest. | node --check tests/rag-retrieval-row-contract.test.ts; git diff --check; ledger-inbox; outstanding-issues; branch-review-ledger; ledger-write-discipline; ci-change-scope --self-test |
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,13 @@
{
"version": 1,
"id": "d194f4ec-568c-4689-a411-22447c59fb53",
"createdOn": "2026-08-15",
"action": "add",
"payload": {
"pri": "P2",
"type": "issue",
"summary": "Recurring 'Unhandled server request error' on /api/search and /api/search/universal is untriaged",
"detail": "Three Sentry issue groups in clinibase-xz over 24h (JAVASCRIPT-NEXTJS-Y, -Z, -10), 17 events, 0 users impacted, all titled 'Error: Unhandled server request error' with culprit chunk 1261.js:2:4801. Top frames are /api/search/route.js and /api/search/universal/route.js. First seen 2026-08-14T08:44:37Z on release c9b089c92c975297c10649b005401d5ae337cf48, roughly six hours BEFORE PR #1946 merged, so it is not caused by the retrieval row contract; the post-merge group is the same error refingerprinted by the release change. The error string does not appear anywhere in repo source, so it likely originates in a dependency or an instrumentation wrapper — origin unidentified. Nobody owns this. Next step: identify what throws it, then decide whether it is a bot/scanner artefact or a real request-handling gap.",
"source": "Sentry clinibase-xz, reviewed 2026-08-15"
}
}
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,13 @@
{
"version": 1,
"id": "fd42e1f9-4012-4296-b7c2-102c7d199738",
"createdOn": "2026-08-15",
"action": "add",
"payload": {
"pri": "P3",
"type": "task",
"summary": "Make the retrieval row contract's source_metadata pin structural, not data-guaranteed",
"detail": "rag-row-contracts.ts pins source_metadata to a JSON object via z.record(...), but documents.metadata is bare jsonb and permits arrays and scalars. Measured against the live project (sjrfecxgysukkwxsowpy) on 2026-08-15: all 2851 documents are object-typed, so nothing breaks today and no live errors exist. The guarantee is data, not schema — a future ingest path could violate it and take retrieval down for that document's chunks. Fix is either a check (jsonb_typeof(metadata) = 'object') constraint on public.documents, or loosening the pin. Every other required field in that contract is backed by a not-null constraint.",
"source": "PR #1946 review + live Supabase verification 2026-08-15"
}
}
68 changes: 35 additions & 33 deletions src/lib/rag/rag-candidate-sources.ts
Original file line numberDiff line numberDiff line change
Expand Up@@ -26,6 +26,14 @@ import {
import { applyMemoryCardBoosts, fetchMemoryCardsForQuery } from "@/lib/deep-memory";
import { env } from "@/lib/env";
import { logger } from "@/lib/logger";
import {
RetrievalRowShapeError,
assertEmbeddingFieldRows,
assertIndexUnitRows,
assertRetrievalRows,
type EmbeddingFieldSignalRow,
type IndexUnitSignalRow,
} from "@/lib/rag/rag-row-contracts";
import {
firstVariantPoolIsStrong,
maxTextRpcQueryVariants,
Expand DownExpand Up@@ -140,6 +148,12 @@ export function recordHybridRpcError(telemetry: SearchTelemetry | undefined, rpc
}
}

/** Index-unit consumers read these provenance fields as maps; retain that output contract. */
function optionalJsonRecord(value: unknown): Record<string, unknown> | null {
if (!value || typeof value !== "object" || Array.isArray(value)) return null;
return value as Record<string, unknown>;
}

/** Record how many variant RPCs a lexical surface actually issued (PT-02 early-exit). */
function recordTextVariantFanout(
telemetry: SearchTelemetry | undefined,
Expand DownExpand Up@@ -211,7 +225,9 @@ export async function searchTextChunkCandidates(args: {
// most-terminal lexical layer surfaces in hybrid_rpc_errors telemetry
// instead of silently degrading to zero candidates. Return value unchanged.
if (error) recordHybridRpcError(args.telemetry, "match_document_chunks_text", error);
return error || !data?.length ? ([] as SearchResult[]) : (data as SearchResult[]);
if (error || !data?.length) return [] as SearchResult[];
assertRetrievalRows(data, "match_document_chunks_text");
return data;
};

const variants = args.queryVariants.slice(0, maxTextRpcQueryVariants);
Expand DownExpand Up@@ -338,14 +354,6 @@ export type ChunkSignalMatch = {
indexUnit?: DocumentIndexUnitMatch | null;
};

type IndexUnitRpcRow = DocumentIndexUnitMatch & {
document_id: string;
source_chunk_id: string | null;
similarity?: number | null;
text_rank?: number | null;
hybrid_score?: number | null;
};

type TableFactRpcRow = {
id: string;
document_id: string;
Expand DownExpand Up@@ -1058,26 +1066,14 @@ export async function searchEmbeddingFieldCandidates(args: {
);
if (error) recordHybridRpcError(args.telemetry, "match_document_embedding_fields_hybrid", error);
if (error || !data?.length) return [] as SearchResult[];
const matches = (
data as Array<{
source_chunk_id: string | null;
field_type: string | null;
similarity?: number | null;
text_rank?: number | null;
hybrid_score?: number | null;
}>
)
.filter(
(
row,
): row is {
source_chunk_id: string;
field_type: string | null;
similarity?: number | null;
text_rank?: number | null;
hybrid_score?: number | null;
} => Boolean(row.source_chunk_id),
)
try {
assertEmbeddingFieldRows(data, "match_document_embedding_fields_hybrid");
} catch (error) {
if (!(error instanceof RetrievalRowShapeError)) throw error;
return [] as SearchResult[];
}
const matches = data
.filter((row): row is EmbeddingFieldSignalRow & { source_chunk_id: string } => Boolean(row.source_chunk_id))
Comment thread
coderabbitai[bot] marked this conversation as resolved.
.map((row) => ({
chunkId: row.source_chunk_id,
similarity: Number(row.similarity ?? 0),
Expand DownExpand Up@@ -1125,8 +1121,14 @@ export async function searchIndexUnitCandidates(args: {
);
if (error) recordHybridRpcError(args.telemetry, "match_document_index_units_hybrid", error);
if (error || !data?.length) return [] as SearchResult[];
const matches = (data as IndexUnitRpcRow[])
.filter((row): row is IndexUnitRpcRow & { source_chunk_id: string } => Boolean(row.source_chunk_id))
try {
assertIndexUnitRows(data, "match_document_index_units_hybrid");
} catch (error) {
if (!(error instanceof RetrievalRowShapeError)) throw error;
return [] as SearchResult[];
}
const matches = data
.filter((row): row is IndexUnitSignalRow & { source_chunk_id: string } => Boolean(row.source_chunk_id))
.map((row) => ({
chunkId: row.source_chunk_id,
similarity: Number(row.similarity ?? 0),
Expand All@@ -1146,13 +1148,13 @@ export async function searchIndexUnitCandidates(args: {
page_end: row.page_end,
heading_path: row.heading_path ?? [],
normalized_terms: row.normalized_terms ?? [],
source_span: row.source_span ?? null,
source_span: optionalJsonRecord(row.source_span),
quality_score: row.quality_score,
extraction_mode: row.extraction_mode,
similarity: row.similarity,
text_rank: row.text_rank,
hybrid_score: row.hybrid_score,
metadata: row.metadata ?? null,
metadata: optionalJsonRecord(row.metadata),
},
}));
return loadChunksForSignalMatches({
Expand Down
80 changes: 76 additions & 4 deletions src/lib/rag/rag-row-contracts.ts
Original file line numberDiff line numberDiff line change
Expand Up@@ -17,9 +17,13 @@ import type { SearchResult } from "@/lib/types";
*
* The schema is deliberately asymmetric:
*
* - **Strict on the ranking, citation, and evidence fields.** The required chunk identity,
* provenance, and visual fields are `not null` in `supabase/schema.sql`, so requiring them
* cannot reject a row that works today. The four score fields are `.nullish()` — absent or
* - **Strict on the ranking, citation, and evidence fields.** Every required field except
* `source_metadata` is `not null` in `supabase/schema.sql`, so requiring it cannot reject a
* row that works today. `source_metadata` is the exception: `documents.metadata` is bare
* `jsonb`, which permits arrays and scalars, so pinning it to an object is guaranteed by the
* data rather than by a constraint. Measured 2026-08-15, all 2851 live documents are
* objects; a `check (jsonb_typeof(metadata) = 'object')` would make that structural.
* The four score fields are `.nullish()` — absent or
* null already flows through the downstream `?? 0` handling unchanged — but a *string where
* a number belongs* is rejected, which is precisely the silent-misranking case this exists
* to catch.
Expand DownExpand Up@@ -100,7 +104,12 @@ function describeIssues(error: z.ZodError): string[] {
* has them, which would otherwise swallow the signal entirely.
*/
export function assertRetrievalRows(rows: unknown, rpc: string): asserts rows is SearchResult[] {
const parsed = retrievalRowsSchema.safeParse(rows);
assertRowsAgainst(retrievalRowsSchema, rows, rpc);
}

/** Shared validate-log-throw step. Kept separate so every row contract fails identically. */
function assertRowsAgainst(schema: z.ZodType, rows: unknown, rpc: string): void {
const parsed = schema.safeParse(rows);
if (parsed.success) return;
const issues = describeIssues(parsed.error);
logger.error("retrieval_row_shape_mismatch", {
Expand All@@ -111,6 +120,69 @@ export function assertRetrievalRows(rows: unknown, rpc: string): asserts rows is
throw new RetrievalRowShapeError(rpc, issues);
}

/**
* Signal rows from `match_document_embedding_fields_hybrid` — not `SearchResult`s.
*
* These carry a chunk id plus scores; `loadChunksForSignalMatches` then loads the real chunk.
* A wrong `source_chunk_id` loads the wrong evidence, and the mapping coerces scores with
* `Number(row.similarity ?? 0)`, which turns a stringified score into a silently different
* number rather than an error. Both are validated here; `field_type` is only a provenance
* label, so it stays permissive.
*/
const embeddingFieldRowSchema = z.looseObject({
source_chunk_id: z.string().nullable(),
field_type: z.string().nullable(),
similarity: z.number().nullish(),
text_rank: z.number().nullish(),
hybrid_score: z.number().nullish(),
});

export type EmbeddingFieldSignalRow = z.infer<typeof embeddingFieldRowSchema>;

/** Validate embedding-field signal rows before they select chunks and scores. */
export function assertEmbeddingFieldRows(rows: unknown, rpc: string): asserts rows is EmbeddingFieldSignalRow[] {
assertRowsAgainst(z.array(embeddingFieldRowSchema), rows, rpc);
}

/**
* Index-unit rows from `match_document_index_units_hybrid`.
*
* The fields that select a chunk, control scoring, or label extraction are backed by
* constraints on `public.document_index_units` in `supabase/schema.sql`: `unit_type`, `title`,
* `content` and `extraction_mode` are `not null`, and both `unit_type` and `extraction_mode`
* carry `check` constraints — so the enum below cannot reject a row the database would accept.
* `source_span` and `metadata` are unconstrained `jsonb`, so they accept any JSON value rather
* than turning a non-object provenance value into a retrieval outage. `heading_path` and
* `normalized_terms` stay `.nullish()` to match the `?? []` handling the caller already applies.
*/
const indexUnitRowSchema = z.looseObject({
id: z.string().min(1),
document_id: z.string().min(1),
source_chunk_id: z.string().nullable(),
source_image_id: z.string().nullable(),
unit_type: z.string().min(1),
title: z.string(),
content: z.string(),
page_start: z.number().int().nullable(),
page_end: z.number().int().nullable(),
heading_path: z.array(z.string()).nullish(),
normalized_terms: z.array(z.string()).nullish(),
source_span: z.json().nullish(),
quality_score: z.number().nullable(),
extraction_mode: z.enum(["deterministic", "model_heavy", "hybrid"]),
metadata: z.json().nullish(),
similarity: z.number().nullish(),
text_rank: z.number().nullish(),
hybrid_score: z.number().nullish(),
});

export type IndexUnitSignalRow = z.infer<typeof indexUnitRowSchema>;

/** Validate index-unit signal rows before they select chunks, scores, and unit provenance. */
export function assertIndexUnitRows(rows: unknown, rpc: string): asserts rows is IndexUnitSignalRow[] {
assertRowsAgainst(z.array(indexUnitRowSchema), rows, rpc);
}

/** Build and validate the locally retrieved rows used as document-summary context. */
export function buildDocumentSummaryResults(
chunks: unknown[],
Expand Down
3 changes: 2 additions & 1 deletion src/lib/rag/rag.ts
Original file line numberDiff line numberDiff line change
Expand Up@@ -2094,7 +2094,8 @@ export async function searchChunksWithTelemetry(
// A1: the embedding-field, index-unit, and chunk-hybrid RPCs each depend only on the
// already-computed query embedding and have no data dependency on one another, so run
// them concurrently instead of as three sequential Supabase round-trips. The two helper
// functions swallow their own RPC errors and resolve to [], so Promise.all cannot reject.
// functions resolve their own RPC errors and logged row-shape mismatches to [], so Promise.all
// cannot reject when an optional signal layer drifts.
throwIfAborted(args.signal);
const parallelRpcStartedAt = Date.now();
const [embeddingFieldResult, indexUnitResult, hybridResult] = await Promise.all([
Expand Down
Loading
Loading