feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543) - #551

Merged
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation
Aug 13, 2026
Merged

feat(quiz): mastery-model seam, honest delivered counts, wire validation, concurrency tests (#543)#551
AndresL230 merged 2 commits into
mainfrom
feat/543-quiz-scoring-generation

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

Closes#543. Workstream E of the pre-revamp quiz repair batch (epic #537).

What

E1 — the mastery model is a seam, not a magic number.MASTERY_DELTA_PER_CORRECT / _PER_WRONG and mastery_after() live in services/quiz_config.py with the pedagogy written down. The numbers do not change: the #393 journey's +0.09 is byte-identical and now pinned by a unit test as well.

The trade-offs the revamp gets to decide are written up in docs/quiz-mastery-model.md: length dominates (a 10-question quiz moves mastery 3.3× a 3-question one, purely for having more items), difficulty is free (the agent varies difficulty by mastery and scoring discards that signal), and adaptive mode is unpriced. Options A–E are laid out with the constraint that any change must update the #393 journey in the same commit.

E2 — generation stops silently short-changing quizzes. The response reports requested_count and delivered_count. Losing more than a third of the requested questions to drift triggers one bounded top-up run (a retry loop against a drifting model burns tokens without converging; a slightly short quiz beats a slow one). A failed top-up serves what we have rather than losing it; all-dropped still 502s.

E3 — wire-format validation at the route boundary. At least two options, no duplicate option text (a student can otherwise pick "the same" answer and be wrong), exactly one correct option (zero is the #129 free point, two means the grader's first match wins silently), and no duplicate question stems within one attempt.

E4 — concurrency tests. Double-answer on one index (the UNIQUE arbitrates; the loser re-reads instead of 500ing) and generate-while-generating for one concept (distinct attempt rows). Double-submit was already pinned by #464's tests.

Note on the rebase

Rebasing onto the merged #550 surfaced a real shadowing bug: D's post-apply_graph_update block assigned a local mastery_after, which shadows E's imported mastery_after() model function — UnboundLocalError on every submit. Caught by 24 existing tests, fixed here (the value lives in mastery_score_after).

Verification

  • Hermetic suite: 1976 passed (16 new tests, written first, watched fail).
  • Full local cycle under the stack lock on the stacked branch: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

…ion, concurrency tests (#543)
Workstream E of the pre-revamp quiz repair batch (epic #537):
E1 — the mastery model is a named seam: services/quiz_config.py holds
MASTERY_DELTA_PER_CORRECT/_PER_WRONG plus mastery_after(), with the
pedagogy written down. THE NUMBERS DO NOT CHANGE — the #393 journey's
+0.09 is byte-identical and pinned by a new test. The options the revamp
gets to choose from (length normalization, difficulty weighting,
diminishing returns) are written up in docs/quiz-mastery-model.md,
including the constraint that any change updates the journey in the same
commit.
E2 — generation stops silently short-changing quizzes: the response
reports requested_count and delivered_count, and losing more than a
third of the requested questions to drift triggers ONE bounded top-up
run (a retry loop against a drifting model burns tokens without
converging). A failed top-up serves what we have; all-dropped still
502s.
E3 — wire-format validation at the route boundary: at least two options,
no duplicate option text (a student can otherwise pick "the same" answer
and be wrong), exactly one correct option, and no duplicate question
stems within one attempt.
E4 — concurrency tests: double-answer on one index (the UNIQUE
arbitrates; the loser re-reads instead of 500ing) and
generate-while-generating for one concept (distinct attempt rows). The
double-submit claim was already pinned by #464's tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:11 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: f705d7c4-62e6-475e-8e11-6438a54c9b5f

📥 Commits

Reviewing files that changed from the base of the PR and between d071770 and a961f14.

📒 Files selected for processing (9)
  • backend/agents/__init__.py
  • backend/routes/quiz.py
  • backend/services/quiz_config.py
  • backend/tests/test_quiz_routes.py
  • backend/tests/test_quiz_scoring_e.py
  • docs/quiz-mastery-model.md
  • frontend/src/components/QuizPanel.test.tsx
  • frontend/src/components/QuizPanel.tsx
  • frontend/src/lib/api.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staginga961f14Commit Preview URL

Branch Preview URL
Aug 13 2026, 05:48 AM

…op over-rejecting, surface short quizzes
The review found the E2 top-up miskeyed at its core, verified by
execution. All ten findings addressed:
- The trigger keyed on requested-minus-delivered, conflating "we rejected
some" with "the agent returned fewer". The Quiz schema lets a run
return any count and the E2E seam always returns 3 against a UI default
of 5, so the top-up fired a second full generation on EVERY quiz
journey — double tokens and ~5s latency for zero extra questions. It
now counts questions actually DROPPED and gates on that.
- The top-up prompt said "different from the ones already asked" without
saying what they were, so a deterministic model re-emitted the same
stems and the dedupe discarded the whole retry. It now lists them.
- `wire_questions and ...` made total drift the ONE case that never
retried — backwards, since that's the case a retry most obviously
clears. Total drift now retries once, then 502s (the old
assert_called_once in test_quiz_routes pinned the wrong behaviour and
is updated with the reasoning).
- The retry reused ORCHESTRATOR_LIMITS, handing it a fresh full budget
and doubling the per-request cost backstop. It gets its own smaller
TOPUP_LIMITS.
- The recovery path logged a traceback on a request that deliberately
succeeds — the exact pattern that reds the logscan oracle. Now a
warning with the exception type/message.
- The duplicate-option check casefolded while grading matches
case-SENSITIVELY, so questions whose distractors differ only by case
(`list` vs `List` — a real question) were dropped. It now compares the
way grading does.
- requested_count/delivered_count had no consumer: QuizPanel now warns
"We could only build N of M questions" instead of quietly serving a
short quiz.
- The double-answer test never reached the race path (its fake
short-circuited on the pre-read); it now models the real interleaving
and asserts the loser actually attempted an insert.
Left as-is with reasoning: two of _validate_wire_question's checks are
unreachable from today's only caller (the agent schema pins 4 options
and the caller builds exactly one correct flag) — they are cheap
defence-in-depth for the #537 revamp's new call sites, and the unit
tests exercise them directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit 150071d into mainAug 13, 2026
6 of 8 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz E: scoring/generation correctness — mastery-delta config seam, honest delivered_count, wire validation, concurrency tests

1 participant

@AndresL230