Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) by AndresL230 · Pull Request #552 · SaplingLearn/Sapling · GitHub
Skip to content

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544) - #552

Merged
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability
Aug 13, 2026
Merged

feat(quiz): rate limit, daily spend guard, generation timeout, failure events (#544)#552
AndresL230 merged 2 commits into
mainfrom
feat/544-quiz-cost-observability

Conversation

@AndresL230

@AndresL230AndresL230 commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Closes#544. Workstream F of the pre-revamp quiz repair batch (epic #537) — the last one.

What

F1 — generation is no longer an unbounded LLM call behind a button.

  • A per-user sliding-window limit (8 per 5 minutes, sized for a human comparing difficulties or retaking a concept — nobody legitimately generates ten in one sitting) returns 429 QUIZ_RATE_LIMITED with a Retry-After header, reusing services/request_limits.py.
  • A daily per-user LLM spend ceiling ($2, read off the llm_usage ledger agents/usage.py::record_agent_usage already writes) returns 429 QUIZ_DAILY_LIMIT_REACHED before the model runs. A default-tier generation costs well under a cent, so this bounds a runaway rather than rationing normal use.

Both guards sit after the ownership check, so probing a stranger's concept can't consume their quota, and the spend check fails open — a usage-table blip must not deny every student.

F2 — explicit timeout. The whole generation (agent run, its tool calls, and E2's top-up) is bounded by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long" rather than the generic failure.

F3 — failures reach admin analytics.quiz.generation_failed (category error, with reason: timeout / agent_guardrail / agent_error) joins the pinned taxonomy, so a 502 the student saw is a 502 an admin can count in the errors feed #548 widened to category-based filtering. A throttled student is deliberately not an error event.

F4 — deep-link scoping tests. A foreign concept 404s before the agent runs; a student's own concept from a past semester still generates. Scoping is by ownership, not active semester — pinned so a future "scope to active semester" change can't silently break revision without a deliberate decision.

Also

The process-global rate-limit state now resets between tests via a conftest autouse fixture, same reasoning as the existing _clear_lru_caches. Without it one test's burst throttles every later test hitting the same route — which is exactly what happened when the limit first landed (24 unrelated failures).

Verification

  • Hermetic suite: 1994 passed (12 new tests, written first, watched fail).
  • Full local cycle under the stack lock: Playwright → oracles → integration → teardown.
  • ruff check clean. No migration in this workstream.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Quiz generation now has per-user rate and daily usage limits.
    • Added clearer handling for timeouts and generation failures, including retry guidance.
    • Pagination support is available for large data queries.
  • Bug Fixes

    • Preserved response headers when handling errors.
    • Partial quiz results can remain available when supplemental generation fails.
    • Rate-limit capacity is restored after failed generations.

…e events (#544)
Workstream F of the pre-revamp quiz repair batch (epic #537):
F1 — generate is no longer an unbounded LLM call behind a button: a
per-user sliding-window limit (8 per 5 minutes, sized for a human
comparing difficulties) returns 429 QUIZ_RATE_LIMITED with Retry-After,
and a daily per-user LLM spend ceiling ($2, read off the llm_usage
ledger agents/usage.py already writes) returns 429
QUIZ_DAILY_LIMIT_REACHED before the model runs. Both guards sit AFTER
the ownership check, so probing a stranger's concept can't consume their
quota, and the spend check fails OPEN — a usage-table blip must not deny
every student.
F2 — the whole generation (agent run + tools + the E2 top-up) is bounded
by QUIZ_GENERATION_TIMEOUT_SEC and maps to its own
QUIZ_GENERATION_TIMEOUT code, so the client can say "that took too long"
instead of the generic failure.
F3 — quiz.generation_failed events (category=error, with a reason:
timeout / agent_guardrail / agent_error) join the pinned taxonomy, so a
502 the student saw is a 502 an admin can count in the errors feed #548
widened. A throttled student is deliberately NOT an error event.
F4 — deep-link scoping tests: a foreign concept 404s before the agent
runs, and a student's OWN concept from a past semester still generates
(scoping is by ownership, not active semester — pinned so a future
"scope to active semester" change can't silently break revision).
Also: the process-global rate-limit state now resets between tests via a
conftest autouse fixture, same reasoning as the lru_cache reset — without
it one test's burst throttles every later test hitting the same route.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Aug 13, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 87e35ce2-1c2b-4b7d-8f85-0cf45d388095

📥 Commits

Reviewing files that changed from the base of the PR and between 150071d and 1485478.

📒 Files selected for processing (12)
  • backend/db/connection.py
  • backend/main.py
  • backend/routes/extract.py
  • backend/routes/gradescope.py
  • backend/routes/quiz.py
  • backend/services/events_service.py
  • backend/services/quiz_config.py
  • backend/services/quiz_errors.py
  • backend/services/request_limits.py
  • backend/tests/conftest.py
  • backend/tests/test_event_capture_seams.py
  • backend/tests/test_quiz_cost_observability_f.py

📝 Walkthrough

Walkthrough

Quiz generation now enforces per-user rate and spend limits, bounds agent runs, refunds failed attempts, and records failure events. HTTP exception handling preserves response headers, rate-limit routes expose Retry-After, and database selects support pagination offsets.

Changes

Quiz generation controls

Layer / File(s)Summary
Pagination and HTTP error contracts
backend/db/connection.py, backend/services/quiz_errors.py, backend/main.py, backend/routes/extract.py, backend/routes/gradescope.py
SupabaseTable.select accepts offset. HTTP exception responses preserve supplied headers. Rate-limit responses include Retry-After. Quiz errors define stable limit and timeout codes and support response headers.
Usage guardrails and generation admission
backend/services/quiz_config.py, backend/services/request_limits.py, backend/routes/quiz.py, backend/tests/conftest.py, backend/tests/test_quiz_cost_observability_f.py
Quiz generation applies an 8-request window and a $2 daily spend cap. Usage reads are paginated and fail open on lookup errors. Ownership, rate-limit, spend-limit, and retry-header behavior are tested.
Bounded generation and failure observability
backend/routes/quiz.py, backend/services/events_service.py, backend/tests/test_event_capture_seams.py, backend/tests/test_quiz_cost_observability_f.py
Agent runs use a 90-second timeout. Failed generations refund rate-limit capacity and emit quiz.generation_failed. Timeout, partial top-up, generic failure, and event behavior are tested.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant QuizRoute as quiz.generate route
participant Limits as request_limits
participant Usage as SupabaseTable
participant Agent as quiz agent
participant Events as events_service
Client->>QuizRoute: Submit quiz generation request
QuizRoute->>Limits: Claim per-user rate-limit slot
QuizRoute->>Usage: Read paginated daily LLM usage
Usage-->>QuizRoute: Return usage rows
QuizRoute->>Agent: Run generation with timeout
Agent-->>QuizRoute: Return questions or failure
QuizRoute->>Limits: Refund slot after generation failure
QuizRoute->>Events: Record quiz.generation_failed
QuizRoute-->>Client: Return quiz or structured error response
Loading

Possibly related PRs

Suggested reviewers:darkest-teddy, jose-gael-cruz-lopez

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/544-quiz-cost-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Aug 13, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Aug 13, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging1485478Commit Preview URL

Branch Preview URL
Aug 13 2026, 06:18 AM

All three F guards failed at what they were added to do, each confirmed
by execution in review:
- The daily spend cap summed an UNPAGED llm_usage select. PostgREST caps
a response at max_rows (1000) and answers 206 — a 2xx — so the sum
plateaued and the ceiling could never trip for the runaway user it
targets. It pages now (with an early exit once the cap is crossed), and
db/connection.py::select gained the `offset` it needed.
- Wrapping the whole generation in asyncio.wait_for raised CancelledError
— a BaseException — straight past #543's serve-what-we-have handler, so
a timed-out top-up threw away a valid partial quiz and returned 502.
Each agent run is bounded individually now, so a top-up timeout is an
ordinary TimeoutError the existing handler degrades from; a completed
3-question generation is served instead of discarded.
- The rate-limit slot was claimed before generation and never refunded,
so eight backend 502s locked a student out for five minutes with a
message saying they'd generated too many quizzes — having received
none. Every failure path now refunds the slot
(services/request_limits.py::refund_rate_limit).
Also from the review:
- The timeout branch caught the builtin OSError-family TimeoutError as
well, relabelling transport socket timeouts as wall-clock timeouts. It
catches only asyncio.TimeoutError now.
- The spend-cap comment claimed cross-feature enforcement it doesn't
provide; it now says what's actually true (the spend measured is
cross-feature, the ceiling is enforced on quiz generation only).
- main.py forwards exc.headers as of this branch, which falsified the
comments in routes/extract.py and routes/gradescope.py explaining why
they couldn't send Retry-After. Both now send it.
Known and accepted: a cancelled agent run never reaches
record_agent_usage, so a timed-out generation's tokens don't land in
llm_usage. Capturing usage from a cancelled pydantic-ai run isn't
available at this seam; the per-run timeout narrows the window
considerably versus cancelling the whole request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 merged commit c622e2f into mainAug 13, 2026
4 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

quiz F: cost, abuse, observability — rate limit + spend guard, timeout/502 taxonomy, failure events, scope tests

1 participant

@AndresL230