feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400) - #443

Merged
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles
Jul 28, 2026
Merged

feat(e2e): deterministic oracle module — python -m e2e_oracles (#400)#443
AndresL230 merged 6 commits into
mainfrom
e2e-chapter2-oracles

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 1 of 3. Closes#400.

Adds backend/e2e_oracles/ — deterministic checks the exploratory-testing harness (#399, next PR) invokes mid-session and post-session. The LLM explores; these oracles decide what counts as a finding.

What's here

  • findings.py — shared Finding record + text/JSON rendering.
  • logscan.py — streaming scan of .e2e/backend.log: 5xx from both the uvicorn access log and the sapling.request correlation lines (aggregated per route), traceback blocks in all three observed shapes (bare, ASGI, ExceptionGroup +/| gutters), ANSI-stripped, with the RAG embedding path sits below the SAPLING_MODEL_MODE seam — live embed calls fire even in function mode #439 RAG-indexing noise allowlisted (suppressed + counted, never findings).
  • judges.py — pure functions: graph payload integrity vs DB (the Duplicate subject_root node in /api/graph → React duplicate-key warnings in KnowledgeGraph2D #355 oracle — duplicate node/edge ids, expected counts = db nodes + distinct enrolled courses / drawable edges + spokes, per-course subject-root multiplicity), ciphertext-at-rest (decrypt-must-succeed-and-differ; injected decrypt_fn), API-vs-DB count integrity, orphaned rows.
  • gather.py + __main__.py — the CLI: venv/bin/python -m e2e_oracles [--json] [--check graph|counts|ciphertext|logscan|orphans] [--user] [--base-url] [--log]. Exit 0 clean / 1 findings / 2 infra error. Raw SQL (psycopg, SELECT-only) is sanctioned here as test tooling, mirroring tests/integration/conftest.py: same exact-hostname loopback guards (raise, never skip) on both SUPABASE_DB_URL and --base-url; session minted via services.session_tokens.
  • 20 hermetic tests (no stack, no DB — the CHECKS registry is monkeypatched).
  • One drive-by: corrected the stale "not encrypted yet" rollout note in services/encryption.pymessages.content / room_messages.text / sessions.summary_json are encrypted (seeder + Chapter 1 journeys prove it at rest).

Live verification (evidence in the SDD report)

Against the seeded stack (function mode, dummy key): full run exits 1 catching the #355 signature — duplicate subject_root__rich-course-cs101, node/edge count mismatches with formula parity to graph.spec.ts's graphExpectations() (absolute numbers reconciled against 2 legitimate non-seed rows in the long-lived local volume); --check ciphertext, orphans, logscan all exit 0 with the logscan result spot-checked against the raw log; --json parses.

Hermetic suite: 1151 passed / 26 skipped (+ the pre-existing #354test_ocr_pipeline loop-affinity error, unrelated). ruff check clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a command-line diagnostic tool for validating end-to-end graph, count, encryption, orphan-record, and backend log health.
    • Added human-readable and JSON reporting with findings, evidence, and suppressed-item counts.
    • Added detection for server errors, tracebacks, data mismatches, missing relationships, and unencrypted values.
    • Restricted diagnostics to local service targets for safer execution.
  • Tests

    • Added comprehensive coverage for CLI behavior, validation checks, log scanning, reporting, and exit codes.
  • Documentation

    • Updated encryption documentation to accurately describe encrypted database fields.

AndresL230and others added 5 commits July 27, 2026 23:36
#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ts, orphans (#400)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review cleanup on #400 judges — collapse the exception and
plain==value paths in ciphertext_findings into one not_encrypted flag;
same Finding either way, less repetition.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Wires the Task 1-2 judges/findings/logscan to the live local stack:
gather.py holds one shared lazily-opened psycopg connection plus authed
httpx calls (mint_session cookie), guarded by require_local (exact-hostname
loopback check) on both SUPABASE_DB_URL and --base-url so this can never
reach staging/production. __main__.py exposes the CHECKS registry and the
python -m e2e_oracles CLI contract (--json/--check/--user/--base-url/--log,
exit 0/1/2) that the explore harness and explorer prompt depend on verbatim.
Hermetic tests monkeypatch CHECKS wholesale -- no DB or HTTP in the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… stale encryption docstring (#400)
- __main__.main() now also sets infra_error when a check RETURNS a
Finding(oracle="oracle-error", ...) (e.g. run_logscan's missing-log-file
case), not just when a check raises. Previously --check logscan against a
nonexistent log exited 1 instead of the contractually-required 2.
- Corrected services/encryption.py's stale ROLLOUT NOTES claim that
messages.content, room_messages.text, and sessions.summary_json were
"not encrypted yet" — all three are encrypted in practice (seeded via
encrypt_if_present/encrypt_json in db/seed_local_rich.py, listed in
CLAUDE.md's Gotchas, and asserted ciphertext-at-rest by
frontend/e2e/tutor.spec.ts and study-room.spec.ts). Comment-only change,
no encryption code or manifest touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:38 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: eb47a006-882d-488d-bc9e-119b309f8a1b

📥 Commits

Reviewing files that changed from the base of the PR and between b7d24dd and e514fa5.

📒 Files selected for processing (2)
  • backend/e2e_oracles/logscan.py
  • backend/tests/test_e2e_oracles_logscan.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-oracles

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment threadbackend/e2e_oracles/logscan.py Fixed
Comment on lines +32 to +33
"2026-07-27 22:17:51,318 [ERROR] routes.documents: [RAG] "
"_index_document_chunks failed for doc 0b65",
Comment on lines +99 to +100
"2026-07-27 22:17:08,060 [INFO] httpx: HTTP Request: POST "
'http://127.0.0.1:54321/storage/v1/bucket "HTTP/1.1 400 Bad Request"',
@@ -0,0 +1,114 @@
"""Hermetic tests for the #400 backend-log scanner. No stack, no DB."""

from e2e_oracles.findings import Finding, render_json, render_text # noqa: F401 — Finding is part of the public interface under test
_DEFAULT_LOG = _REPO_ROOT / ".e2e" / "backend.log"

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),

CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
CHECKS: dict[str, Callable[[argparse.Namespace], tuple[list[Finding], int]]] = {
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"graph": lambda args: gather.run_graph(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"counts": lambda args: gather.run_counts(args),
"ciphertext": lambda args: gather.run_ciphertext(args),
"logscan": lambda args: gather.run_logscan(args),
"orphans": lambda args: gather.run_orphans(args),
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginge514fa5Commit Preview URL

Branch Preview URL
Jul 28 2026, 07:53 AM

@coderabbitai

Copy link
Copy Markdown

Caution

Failed to replace (edit) comment. This is likely due to insufficient permissions or the comment being deleted.

Error details
putComment timed out

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 2 issues:

  1. KEY_RE never matches a bare Exception: msg / Error: msg terminal line — \w+ must consume at least one word char before the Error|Exception suffix, so for Exception: boom it eats the whole word and the suffix match fails. _block_key then falls back to block_lines[0], which is the generic Traceback (most recent call last): for every plain traceback, so two unrelated tracebacks that end in bare Exception: lines aggregate into a single "1 traceback (2×)" finding, hiding a distinct bug. (Verified by running the regex: ValueError: boom matches, Exception: boom does not.)

# or " | KeyError: 'z'" (ExceptionGroup gutter prefix).
KEY_RE=re.compile(r"^[\s|+]*\w+(\.\w+)*(Error|Exception|Interrupt|Exit)\b.*")

  1. NEW_LOG_LINE_RE doesn't recognize the fourth documented line format — the ANSI pre-request lines (22:17:08.488 GET /api/health after stripping) match neither the INFO:-style nor the YYYY-MM-DD alternative, and while a traceback block is open only NEW_LOG_LINE_RE closes it. Any traceback immediately followed by such a line swallows it into the block's excerpt, and a second traceback following without an intervening INFO:/date-prefixed line merges into the first block entirely. The module's own header lists this format as one the scanner must handle.

# A line that unambiguously starts a *new* log record — closes an open block.
NEW_LOG_LINE_RE=re.compile(r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} ")

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

…e-request block close (PR #443)
Two PR-review findings on backend/e2e_oracles/logscan.py:
- KEY_RE required a class-name prefix (e.g. `ValueError`) before
Error/Exception/Interrupt/Exit, so bare `Exception: boom` terminal
lines never matched; _block_key then fell back to block_lines[0],
aggregating unrelated tracebacks under the generic "Traceback (most
recent call last):" key. Made the prefix optional.
- NEW_LOG_LINE_RE didn't recognize the post-ANSI-strip HH:MM:SS.mmm
pre-request log line shape, so it got swallowed into an open
traceback block instead of closing it, merging a following
traceback into the same block. Added that alternative.
Added three regression tests (RED before fix, GREEN after); the
existing ExceptionGroup test still resolves to the KeyError line
since it's still the last KEY_RE match in that block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# The third alternative is the post-ANSI-strip shape of the pre-request log line
# (e.g. "22:17:08.488 GET /api/health") — scanning runs after ANSI stripping.
NEW_LOG_LINE_RE = re.compile(
r"^(INFO|ERROR|WARNING|DEBUG|CRITICAL):|^\d{4}-\d{2}-\d{2} |^\d{2}:\d{2}:\d{2}\.\d+ "
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Both review findings fixed in e514fa5 (bare-Exception KEY_RE match + ANSI pre-request block close), TDD'd with 3 new regression tests; scoped re-review confirms no regex regressions (ExceptionGroup keying, ASGI header, chained-exception blocks all probed).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(agent): invariant oracles for unscripted exploration

1 participant

@AndresL230