fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(rag): put retrieval/indexing failures on the app-logging path (#482) - #501

Merged
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability
Jul 31, 2026
Merged

fix(rag): put retrieval/indexing failures on the app-logging path (#482)#501
AndresL230 merged 2 commits into
mainfrom
fix/482-rag-observability

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Part of #482.

Scope, after re-verifying the issue against main

The issue lists three hazards. Only one and a half were still live:

#ClaimState
1Silent-empty: bare print, invisible to app logginglive — fixed here
2Poisoned embedding: None rows persisted anywayalready fixed on main (rag_service.py filters them before upsert)
3Durability inversionlive — documented here, code fix deliberately deferred

Item 1

rag_service reported every failure with a bare print, so failures never reached the app-logging path and no rollup could count them. That matters most for retrieve_chunks, which degrades to [] — and [] is also what "nothing relevant matched" returns, so an ungrounded tutor/quiz turn was indistinguishable from a grounded one at every layer above.

Adds a module logger, moves all three print sites onto it, and emits rag.retrieval_failed / rag.chunks_dropped (category error) so the degrade rates are countable.

The part worth reviewing

The naive version of this is wrong, and the existing #439 tests caught it. _require_real_mode() raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != "real" — which is every function-mode run. So a blanket except Exception would emit an error event on every e2e tutor turn and every e2e upload, and the new signal would be pure noise the day it shipped.

_EmbeddingDisabled is now caught separately: INFO, no event. Only genuine failures warn and spend an event. Both #439 transport guards are updated to assert the new channel and to pin the no-event contract — the invariant they protect (no egress in function mode) is unchanged; only the output moved off stdout.

Item 3 — documented, not fixed

Recorded in docs/architecture.md under Known sharp edges: DBOS wraps /upload/sync, which never indexes; the route that does index is the streaming /upload, which is non-durable by design (ADR 0011). A crash between _persist_document and the post-roll indexing task leaves a healthy-looking documents row whose content is absent from retrieval permanently, and scripts/backfill_document_chunks.py is invoked by nothing.

Wiring a retry trigger is a real design call (where it runs, what re-drives it, how it interacts with the streaming route's X-Request-ID replay semantic), so I documented the gap rather than guessing at it — the issue itself asks for documentation as the minimum bar.

Gates

Backend pytest 1531 passed / 32 skipped (2 new), ruff clean, full local e2e cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Item 1 of #482. rag_service reported every failure with a bare `print`, so a
retrieval that blew up was invisible to app logging and uncountable by any
rollup — and since retrieve_chunks degrades to [], which is also what "nothing
relevant matched" returns, an ungrounded tutor/quiz turn was indistinguishable
from a grounded one at every layer above it.
Adds a module logger, moves all three print sites onto it, and emits
rag.retrieval_failed / rag.chunks_dropped so the degrade rates are countable.
Crucially these separate the DELIBERATE degrade from a real one. The #439 seam
raises _EmbeddingDisabled whenever SAPLING_MODEL_MODE != real, which is every
function-mode run — so a naive version emitted an error event on every e2e
tutor turn and every e2e upload. That noise would have made the new signal
worthless. _EmbeddingDisabled now logs at INFO and spends no event; only
genuine failures warn and count. The two #439 transport guards caught this and
are updated to assert the new channel (and to pin the no-event contract) — the
invariant they protect is unchanged, only the output moved off stdout.
Item 2 (poisoned embedding:None rows) was already fixed on main — the filter
at the upsert drops them. Item 3 (durability inversion) is documented in
architecture.md: indexing rides the NON-durable streaming route while DBOS
wraps the sync route that never indexes, so a crash loses chunks behind a
healthy documents row, and backfill_document_chunks.py is invoked by nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc03c983Commit Preview URL

Branch Preview URL
Jul 31 2026, 06:17 PM

@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d39521e0-b7c4-44b7-80b5-a3f250b372fa

📥 Commits

Reviewing files that changed from the base of the PR and between fa111a1 and c03c983.

📒 Files selected for processing (5)
  • backend/services/events_service.py
  • backend/services/rag_service.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_rag_service.py
  • docs/architecture.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s mode
Code review, two real findings.
1. The two new event types were emitted OUTSIDE the pinned #117 taxonomy.
log_event deliberately doesn't enforce membership (it must never raise), so
nothing failed — they were simply undocumented and absent from the frozenset
that exists to make a rename break loudly. Registered in EVENT_TAXONOMY, the
docstring table, and the exact-set test.
Also recorded why they are NOT named error.*: /api/admin/analytics/errors
filters on `event_type like error.*` and projects an HTTP shape (path /
method / status_code / duration_ms). Renaming would fill an HTTP-request
table with null-path rows; these surface via /usage/summary's by_event_type
instead. Giving that feed a shape-agnostic projection is the real fix and is
out of scope here.
2. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient
SAPLING_MODEL_MODE. _require_real_mode() runs before the mocked client is
reached, so with function mode exported — which the E2E workflow tells you
to export — it took the _EmbeddingDisabled branch and passed vacuously on an
unrelated path. Pinned to real mode; verified it now passes with
SAPLING_MODEL_MODE=function set AND unset.
Six PRE-EXISTING tests in that file share the same latent dependence
(task_type, returns_count, chunk_ids, dedupe). Left alone — the hermetic
lane runs with the var unset by contract — but worth its own cleanup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues, both fixed in c03c983.

  1. The two new event types were emitted outside the pinned [P2] Observability: instrument capture seams (middleware, auth, feature routes) #117 taxonomy. log_event deliberately does not enforce membership (it must never raise), so nothing failed — they were simply undocumented and missing from the frozenset whose whole job is to make a rename break loudly.

https://github.com/SaplingLearn/Sapling/blob/5d741c9/backend/services/rag_service.py#L148-L158

Registered in EVENT_TAXONOMY, the docstring table, and the exact-set test. Also recorded why they are not named error.*: /api/admin/analytics/errors filters event_type like error.* and projects an HTTP shape (path/method/status_code/duration_ms), so renaming would fill an HTTP-request table with null-path rows. These surface via /usage/summary's by_event_type instead. Giving that feed a shape-agnostic projection is the real fix, and is out of scope here.

  1. test_retrieve_chunks_failure_is_logged_and_counted depended on ambient SAPLING_MODEL_MODE (bug due to _require_real_mode() running before the mocked _client is reached). With function mode exported — which the E2E workflow instructs — it took the _EmbeddingDisabled branch and passed vacuously on an unrelated code path. Reproduced, then pinned to real mode and verified passing with the var both set and unset.

https://github.com/SaplingLearn/Sapling/blob/c03c983/backend/tests/test_rag_service.py#L389-L400

Noted, not fixed: six pre-existing tests in that file share the same latent dependence (task_type, returns_count, chunk_ids, dedupe). The hermetic lane runs with the var unset by contract, so they are green in CI, but the file deserves its own cleanup.

Checked and cleared: _EmbeddingDisabled is the only type _require_real_mode raises and is caught ahead of the generic handler in both call sites; the embedding_disabled batch flag is correct given model_mode() cannot change mid-call; normal empty results never emit an event; log_event is non-blocking and cannot raise into the request path; no import cycle (the keyless import main guard still passes).

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit 729a6ff into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/482-rag-observability branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230