fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(e2e): poll for the tutor turn's rows instead of reading once (#477) - #500

Merged
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll
Jul 31, 2026
Merged

fix(e2e): poll for the tutor turn's rows instead of reading once (#477)#500
AndresL230 merged 2 commits into
mainfrom
fix/477-tutor-persistence-poll

Conversation

@AndresL230

@AndresL230AndresL230 commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Part of #477.

Cause

The persistence assert read messages once, immediately after the reply text rendered — treating "the reply is on screen" as "the rows are committed". Those are different signals.

The server side is correct: stream_agent_turn persists inside on_complete and only then yields done (backend/services/chat_stream.py:372-393), so the write does precede the end of the turn. But the composer renders off the streamed tokens, so both UI assertions can pass while done is still in flight — leaving a window where the one-shot read sees only the 4 seeded rows. That matches the reported failure exactly (expected 6, received 4, UI assertions green).

So this is a test-side assumption, not a product regression — worth stating, since the alternative reading (rows should already be there) would have pointed at a real persistence bug.

Fix

Bounded expect.poll on the row count, the idiom events.spec.ts already uses for its fire-and-forget rollup. Root fix rather than a retry, per the #388 zero-flake policy.

Honest limitation

A ~1-in-6 race is not reproducible on demand, so this is not verified by a watched-red test. It rests on the persist-before-done contract above plus green cycles. What the change does guarantee: the assert can no longer fail because it read too early — it now waits for the condition it depends on, or fails after 5s.

Gates

tsc clean; full local cycle green (Playwright 37/37, oracles 0 findings).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Improved tutor end-to-end test reliability by waiting for expected messages to be persisted before validation.
    • Updated assertions to verify message roles and encrypted content more consistently.

The persistence assert read `messages` a single time, immediately after the
reply text rendered. That treats "the reply is on screen" as "the rows are
committed", and they are not the same signal.
The server side is fine: stream_agent_turn persists inside on_complete and
only THEN yields `done` (services/chat_stream.py:372-393), so the write does
precede the end of the turn. But the composer renders off the streamed
TOKENS, so both UI assertions can pass while `done` is still in flight —
leaving a window where the one-shot read sees only the 4 seeded rows. That is
exactly the observed failure (expected 6, received 4, UI assertions green).
Replaces the read with a bounded expect.poll on the row count, the idiom
events.spec.ts already uses for its fire-and-forget rollup. Root fix rather
than a retry, per the #388 zero-flake policy.
Note: a ~1-in-6 race is not reproducible on demand, so this is verified by the
persist-before-done contract above plus repeated green cycles, not by a
watched-red test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 31, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging749c76eCommit Preview URL

Branch Preview URL
Jul 31 2026, 05:55 PM

@coderabbitai

coderabbitaiBot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:51 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: bbeabad4-aa38-4b46-882b-3b5e6625fc11

📥 Commits

Reviewing files that changed from the base of the PR and between a621382 and 749c76e.

📒 Files selected for processing (1)
  • frontend/e2e/tutor.spec.ts
📝 Walkthrough

Walkthrough

The tutor E2E test now polls the database for the expected persisted user and assistant messages before validating their roles and encrypted contents.

Changes

Tutor persistence validation

Layer / File(s)Summary
Poll persisted tutor rows
frontend/e2e/tutor.spec.ts
The test replaces the immediate database read with a bounded Playwright poll. It waits for the seeded baseline and two newly persisted message rows before validating the results.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly and concisely describes the main change: polling for tutor message rows instead of reading them once.
Description check✅ PassedThe description clearly explains the cause, fix, limitation, testing, and related issue, although it does not follow every template heading.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/477-tutor-persistence-poll

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Code review: polling the count and then re-querying for the content
assertions decoupled two checks the original single read had joined. A
duplicate-persistence regression could land a row in that gap and still slice
two valid-looking rows off the end — exactly what this journey exists to
catch. Capture inside the predicate, matching events.spec.ts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Code review

Found 1 issue, fixed in 749c76e.

  1. Polling the count and then re-querying for the content assertions decoupled two checks the original single read had joined (bug due to const rows = await readTurnRows() running as a separate query after the poll).

A duplicate-persistence regression could land a row in that gap; the poll's exact toBe(N) would already have passed against the earlier read, and the content assertions would still slice two valid-looking rows off the end and pass. That is precisely the invariant this journey exists to prove (#397 posture). frontend/e2e/events.spec.ts avoids this by capturing the payload as a side effect inside the poll predicate — same shape now used here.

https://github.com/SaplingLearn/Sapling/blob/749c76e/frontend/e2e/tutor.spec.ts#L88-L101

Checked and cleared: the poll is not weaker than the original assert (same exact toBe, just retried); the persist-before-done claim holds (on_complete at chat_stream.py:372, done yielded at :393, and _persist in learn.py is a synchronous save_message); no CLAUDE.md violation.

🤖 Generated with Claude Code

@AndresL230
AndresL230 merged commit fa111a1 into mainJul 31, 2026
7 checks passed
@AndresL230
AndresL230 deleted the fix/477-tutor-persistence-poll branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@AndresL230