ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459

Merged
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps
Jul 29, 2026
Merged

ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main#459
AndresL230 merged 1 commit into
mainfrom
ci/e2e-lock-deps

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Summary

Every e2e.yml run on main has been red since the #349 merge. Root cause: the e2e lane installs requirements.txt, whose unpinned pydantic-ai-slim[google]>=0.0.20 now freshly resolves to the 2.x major — where run_stream_events() returns an async context manager, so chat_stream.py's async for raises TypeError before the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows. ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).

Reproduced tonight by driving the real chat_tutor agent through stream_agent_turn on 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.

Changes

  • e2e.yml: install --require-hashes -r requirements.lock (cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.
  • requirements.txt: pydantic-ai-slim[google]>=0.0.20,<2 with a comment tying the ceiling to the chat_stream.py migration it requires.

Verification

  • e2e.yml is main-only, so the proof lands on the merge commit's run.
  • Locally: full app boot + the exact tutor-journey /chat/stream request on the locked versions → 200, streamed constant, exactly 2 message rows.

🤖 Generated with Claude Code

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 29, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingb57574bCommit Preview URL

Branch Preview URL
Jul 29 2026, 12:22 PM

…c-ai 2.x broke every main e2e run since #349
The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major
on every run; 2.x's run_stream_events() returns an async context
manager, so async-for raises TypeError pre-token, every streamed tutor
turn fell to the legacy fallback (orphan user row + dummy-key failure),
and the tutor journey failed 7-rows-vs-6 on every push to main — while
ci.yml's lock-pinned backend lane (1.107) stayed green.
- e2e.yml now installs --require-hashes -r requirements.lock (pip cache
keyed on the lock), matching ci.yml.
- requirements.txt pins pydantic-ai-slim <2; bump only with a
chat_stream.py migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 29, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:55 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 004dd9ff-d227-4c69-aa1a-88efb0d1bb36

📥 Commits

Reviewing files that changed from the base of the PR and between d0d8837 and b57574b.

📒 Files selected for processing (2)
  • .github/workflows/e2e.yml
  • backend/requirements.txt
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Fix failing CI checks
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/e2e-lock-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@AndresL230
AndresL230 merged commit 7fa6cb6 into mainJul 29, 2026
6 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…s (#151a, 1/2) (#472)
* refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2)
Part one of the final gemini_service cutover (#151): everything learn.py/
streaming. Part two (documents.py legacy pipelines, the file deletion,
ADR 0024) follows; the issue closes with it.
- stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback,
SAME contract (fallback owns persistence + usage; at-most-one-of with
on_complete; error rungs run neither). Rung 1 now degrades to a fresh
NON-STREAMING agent turn on the fast tier (a different, faster model is
a materially better second chance than the same one re-streamed), wired
through the extracted _chat_turn_json / _start_session_agent.
- The writes-guard generalized (#470's blank-reply rule → ALL fallback
entries): if tools already wrote graph/mastery, no fallback ever runs —
terminal error with the new additive retryable:false field. The client
honors it (and 413s): ChatStreamError.retryable +
shouldFallBackToJson(), so Learn's ladder can no longer silently re-run
a turn whose side effects landed (the pre-existing hole that defeated
#470's server guard from the client side).
- Guardrail → status mapping on /chat, /start-session, /action (the notes
precedent): UsageLimitExceeded → 413 naming the cause (deterministic —
the client does NOT retry it), UnexpectedModelBehavior → 502
retry-friendly, bare Exception → 502 + exception log.
- /start-session's JSON route gets its FIRST agent implementation
(_start_session_agent; the legacy pipeline was its primary, not a
fallback), converging the greeting prompt on what /start-session/stream
already shipped. /action agent-ified in place (assistant-only persist
preserved; task-dispatch means the existing chat_tutor handler covers
both — pinned by a new function-mode route test).
- Deleted: _legacy_chat, build_system_prompt, get_conversation_history,
_get_course_documents, _resolve_legacy_model, the template loader, the
five legacy prompt files (grep-verified single reader), and
compact_graph_context. chat.message_sent now has exactly one JSON-path
emission site (inside _chat_turn_json).
- test_streaming_rung1_live.py redesigned: broken-model streaming agent +
good fast-tier Agent.run() fallback — still proving the cross-version
exception-wrapping seam (#459's failure class) live.
- New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO
journeys touched /start-session): entry screen → deterministic greeting
→ lazy-session contract (no row until the first follow-up) → DB-polled
transcript. New testids registered in docs/frontend-testids.md.
Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all
touched stream/agent/route files; frontend 350 passed + tsc clean; evals
replay green ×6 (prompts untouched by design).
Part of #151 (do not auto-close).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended
- The writes-guard now reaches INSIDE the fallback: _chat_via_agent and
_start_session_agent stamp sapling_wrote (their own deps' write-state)
on any post-run exception, and _rung1_fallback_events reads it — a
fallback that wrote graph/mastery and then failed emits retryable:false
so the client cannot re-run the turn a third time and re-apply the
writes (the double-apply class, one level deeper than #470's guard).
Red-first stream tests (wrote-then-failed → not retryable; clean
failure → retryable) + stamp tests at the helper level.
- ADR 0020's 'Retry is already safe' argument amended: transcript
persistence is still exactly-once, but tool writes can land mid-turn —
retryable:false / 413 gate the automatic re-runs now.
Backend 1515 + ruff green; lockvenv 77 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the ci/e2e-lock-deps branch August 2, 2026 18:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230