docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401) - #445

Merged
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook
Jul 28, 2026
Merged

docs(e2e): Chapter 2 exploration runbook — kick-off, triage, promotion (#401)#445
AndresL230 merged 5 commits into
mainfrom
e2e-chapter2-runbook

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Chapter 2 of the E2E program (epic #403), PR 3 of 3. Closes#401.

The operator's manual for make explore / /explore (PR #444) and the oracle CLI (PR #443): what Chapter 2 is (vs Chapter 1's scripted journeys — local-only, never gates), prerequisites, both kick-off modes with the full knob table, the .explore/ output inventory, standalone oracle usage, the triage protocol for findings.md (reproduce → issue with evidence; noise → tune the oracle; known-bug → evidence only if new), the promotion pipeline (issue → scripted Chapter 1 journey, testids per docs/frontend-testids.md; epic bar: ≥3 findings promoted into Chapter 1 regressions in the first month), cadence/cost, the machine-singleton lock protocol, and troubleshooting.

Written against the MERGED harness (post-review reality): unconditional dummy GEMINI_API_KEY + SEED_RICH=1, Edit(.explore/**)-scoped explorer writes, EXPLORE_USER five-user wiring, holder-written lock.ok bookkeeping, session recordings under .explore/traces/. Every command, knob, default, path, and exit code in the doc was verified against the code (knob table cross-checked with /usr/bin/grep -o 'EXPLORE_[A-Z_]*', check names against e2e_oracles.__main__.CHECKS, make -n targets).

Plus one pointer line in docs/local-supabase.md's "One-command E2E stack" section.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Added a comprehensive guide for AI-driven exploratory end-to-end testing.
    • Documented headless and interactive test runs, prerequisites, generated artifacts, oracle usage, troubleshooting, and issue triage.
    • Added guidance for promoting exploratory findings into regression coverage.
    • Linked the new exploratory testing guide from the local Supabase documentation.

AndresL230and others added 3 commits July 28, 2026 03:28
…n pipeline (#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Self-review catch: task-8-report.md lives under .superpowers/ (gitignored),
so pointing the shipped runbook at it as evidence was a dead reference for
anyone without that local planning history. Describe the acceptance-testing
evidence inline instead, and fix "earlier round of the same run" to the
accurate "separate run of the same acceptance round" (task-8-report.md's
run 4 found the tutor-resume 500; run 5 found the wrong-data bug).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 28, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingdf2672bCommit Preview URL

Branch Preview URL
Jul 28 2026, 10:57 AM

@coderabbitai

coderabbitaiBot commented Jul 28, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 202a3c9c-88e8-431a-a2d8-e3646f439c01

📥 Commits

Reviewing files that changed from the base of the PR and between 9f64d66 and df2672b.

📒 Files selected for processing (2)
  • docs/e2e-exploration.md
  • docs/local-supabase.md

📝 Walkthrough

Walkthrough

Adds a complete Chapter 2 exploratory-testing runbook covering execution modes, prerequisites, artifacts, oracle validation, finding triage, regression promotion, locking, cadence, and troubleshooting. Links the runbook from the local Supabase documentation.

Changes

Exploratory testing runbook

Layer / File(s)Summary
Exploration execution and outputs
docs/e2e-exploration.md
Documents prerequisites, headless and interactive workflows, generated artifacts, and stack startup and teardown behavior.
Oracle validation and result interpretation
docs/e2e-exploration.md
Describes standalone oracle commands, selective checks, seeded users, stack assumptions, and exit codes.
Finding triage and regression promotion
docs/e2e-exploration.md
Defines reproduction-based triage and promotion of confirmed findings into scripted Chapter 1 coverage.
Operational procedures and documentation entry point
docs/e2e-exploration.md, docs/local-supabase.md
Adds cadence, cost, locking, troubleshooting, and a local Supabase link to the runbook.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

  • #403 — The runbook documents the Chapter 2 AI Agent Explorer workflow, oracle pipeline, and finding-promotion process described by this epic.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch e2e-chapter2-runbook

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue:

  1. The triage section's flagship example presents a "wrong-data" upload bug (a newly uploaded document coming back with another document's summary/category/concepts) as a genuine promotable finding — but Chapter 2 always runs under SAPLING_MODEL_MODE=function with agents.function_handlers_e2e, whose classifier/summary/concepts handlers return the identical canned output ("Gradient Descent" abstract, lecture_notes category) for every upload by documented design ("Fixed output. Replies are constants…"). The observed symptom is the deterministic seam working as designed — precisely the "noise the harness produced" category the doc's own triage protocol says to drop — not a data-leak bug. Enshrining it as the worked example will teach operators to misfile seam artifacts. (The tutor-resume 500 from the same run — concept_describe unregistered — remains a genuine example and a better candidate for this slot.)

Real evidence for what a genuine finding looks like: this harness's own
bounded acceptance testing turned up a novel "wrong-data" bug (a newly
uploaded, unrelated document came back with another document's summary,
category, and extracted concepts verbatim — polluting the graph with wrong
concepts) and, in a separate run of the same acceptance round, a real
tutor-resume `500` on resuming a session (`POST
/api/graph/<user>/concept-description` failing with `LookupError: ... no
handler is registered for task 'concept_describe'``concept_describe` is
genuinely unregistered in `backend/agents/function_handlers_e2e.py`),
independently corroborated by the `logscan` oracle picking up the same `500`
and traceback. Neither of those is on the known-bugs list — they're exactly
the kind of thing this chapter is for. By contrast, #355 (the graph's
duplicated CS subject-root hub) shows up in the graph oracle on essentially
every run against the rich seed data — that's the known-not-new case
category 3 above exists for.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

AndresL230and others added 2 commits July 28, 2026 03:49
…iming, cross-refs (#401)
Code-review on PR #445 found the flagship "wrong-data upload" example was a
seam artifact (function-mode's E2E_DOC_* fixtures return identical canned
output for every upload by design), not a real bug — reframe it as a worked
"drop + improve the harness" triage example and promote the genuinely novel
tutor-resume 500 (concept_describe unregistered in
agents/function_handlers_e2e.py) to the flagship promotable example instead.
Also: reconcile the "few minutes" (§3) vs "~10 minutes" (§9) timing claims
into one warm-cache-vs-first-run story, and drop two dangling section
cross-refs (§6's inline traces/ caveat didn't need a pointer; "forces both
(see §3)" pointed at a section that never explained the fact) plus align the
--check example with the oracle's own sorted default order.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#401)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review finding fixed in 56511af: the promotable worked example is now the genuine tutor-resume 500 (concept_describe unregistered — facts re-verified against _providers.py/graph.py/main.py by the re-review), and the wrong-data episode is recast as a seam-artifact triage lesson with the recognition tell. Timing contradiction and dangling cross-refs fixed in the same commit; a code-span rewrap nit fixed in the follow-up. Scoped re-review: all findings addressed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly agent run + finding-promotion pipeline (non-gating)

1 participant

@AndresL230