Uh oh!
There was an error while loading. Please reload this page.
ci(e2e): pin the e2e lane to requirements.lock — pydantic-ai 2.x broke the streamed tutor on main - #459
Conversation
Deploying with |
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs | frontend-staging | b57574b | Commit Preview URL Branch Preview URL | Jul 29 2026, 12:22 PM |
…c-ai 2.x broke every main e2e run since #349 The e2e lane freshly resolved pydantic-ai-slim>=0.0.20 to the 2.x major on every run; 2.x's run_stream_events() returns an async context manager, so async-for raises TypeError pre-token, every streamed tutor turn fell to the legacy fallback (orphan user row + dummy-key failure), and the tutor journey failed 7-rows-vs-6 on every push to main — while ci.yml's lock-pinned backend lane (1.107) stayed green. - e2e.yml now installs --require-hashes -r requirements.lock (pip cache keyed on the lock), matching ci.yml. - requirements.txt pins pydantic-ai-slim <2; bump only with a chat_stream.py migration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
18a963c to
b57574bCompareWarning Review limit reached
Next review available in:55 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Uh oh!
There was an error while loading. Please reload this page.
…s (#151a, 1/2) (#472) * refactor(learn): agent-only rung ladder — retire the legacy chat paths (#151a, part 1 of 2) Part one of the final gemini_service cutover (#151): everything learn.py/ streaming. Part two (documents.py legacy pipelines, the file deletion, ADR 0024) follows; the issue closes with it. - stream_agent_turn's seam renamed legacy_fallback → nonstream_fallback, SAME contract (fallback owns persistence + usage; at-most-one-of with on_complete; error rungs run neither). Rung 1 now degrades to a fresh NON-STREAMING agent turn on the fast tier (a different, faster model is a materially better second chance than the same one re-streamed), wired through the extracted _chat_turn_json / _start_session_agent. - The writes-guard generalized (#470's blank-reply rule → ALL fallback entries): if tools already wrote graph/mastery, no fallback ever runs — terminal error with the new additive retryable:false field. The client honors it (and 413s): ChatStreamError.retryable + shouldFallBackToJson(), so Learn's ladder can no longer silently re-run a turn whose side effects landed (the pre-existing hole that defeated #470's server guard from the client side). - Guardrail → status mapping on /chat, /start-session, /action (the notes precedent): UsageLimitExceeded → 413 naming the cause (deterministic — the client does NOT retry it), UnexpectedModelBehavior → 502 retry-friendly, bare Exception → 502 + exception log. - /start-session's JSON route gets its FIRST agent implementation (_start_session_agent; the legacy pipeline was its primary, not a fallback), converging the greeting prompt on what /start-session/stream already shipped. /action agent-ified in place (assistant-only persist preserved; task-dispatch means the existing chat_tutor handler covers both — pinned by a new function-mode route test). - Deleted: _legacy_chat, build_system_prompt, get_conversation_history, _get_course_documents, _resolve_legacy_model, the template loader, the five legacy prompt files (grep-verified single reader), and compact_graph_context. chat.message_sent now has exactly one JSON-path emission site (inside _chat_turn_json). - test_streaming_rung1_live.py redesigned: broken-model streaming agent + good fast-tier Agent.run() fallback — still proving the cross-version exception-wrapping seam (#459's failure class) live. - New greeting-turn journey in tutor.spec.ts (the scoping pass found ZERO journeys touched /start-session): entry screen → deterministic greeting → lazy-session contract (no row until the first follow-up) → DB-polled transcript. New testids registered in docs/frontend-testids.md. Gates: backend 1511 passed + ruff clean; lockvenv 192 passed across all touched stream/agent/route files; frontend 350 passed + tsc clean; evals replay green ×6 (prompts untouched by design). Part of #151 (do not auto-close). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(#472): fix both findings — fallback write-state surfaces to retryable; ADR 0020 amended - The writes-guard now reaches INSIDE the fallback: _chat_via_agent and _start_session_agent stamp sapling_wrote (their own deps' write-state) on any post-run exception, and _rung1_fallback_events reads it — a fallback that wrote graph/mastery and then failed emits retryable:false so the client cannot re-run the turn a third time and re-apply the writes (the double-apply class, one level deeper than #470's guard). Red-first stream tests (wrote-then-failed → not retryable; clean failure → retryable) + stamp tests at the helper level. - ADR 0020's 'Retry is already safe' argument amended: transcript persistence is still exactly-once, but tool writes can land mid-turn — retryable:false / 413 gate the automatic re-runs now. Backend 1515 + ruff green; lockvenv 77 green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
Every
e2e.ymlrun on main has been red since the #349 merge. Root cause: the e2e lane installsrequirements.txt, whose unpinnedpydantic-ai-slim[google]>=0.0.20now freshly resolves to the 2.x major — whererun_stream_events()returns an async context manager, sochat_stream.py'sasync forraisesTypeErrorbefore the first token. Every streamed tutor turn silently fell to the Rung-1 legacy fallback (which writes the user row, then dies on the CI dummy Gemini key), the client then succeeded via the JSON route, and the tutor journey failed its 6-row readback with 7 rows.ci.yml's backend lane stayed green the whole time because it installs the hash-pinned lock (1.107).Reproduced tonight by driving the real chat_tutor agent through
stream_agent_turnon 2.20.0 (TypeError: 'async for' requires an object with __aiter__ method, got _RunStreamEventsContext), and verified the full app + journey request works on the locked set.Changes
e2e.yml: install--require-hashes -r requirements.lock(cache keyed on the lock), matching ci.yml — one dependency universe across CI lanes.requirements.txt:pydantic-ai-slim[google]>=0.0.20,<2with a comment tying the ceiling to thechat_stream.pymigration it requires.Verification
/chat/streamrequest on the locked versions → 200, streamed constant, exactly 2 message rows.🤖 Generated with Claude Code