Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite by Jose-Gael-Cruz-Lopez · Pull Request #80 · SaplingLearn/Sapling · GitHub
Skip to content

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite - #80

Merged
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap
May 5, 2026
Merged

perf(tutor): cap Pro thinking budget + switch Fast to Flash-Lite#80
Jose-Gael-Cruz-Lopez merged 3 commits into
mainfrom
perf/tutor-thinking-cap

Conversation

@Jose-Gael-Cruz-Lopez

Copy link
Copy Markdown
Member

Summary

  • Smart faster: cap gemini-2.5-pro thinking budget at 2048 tokens (was dynamic / unbounded). Applies on both the legacy call_gemini_multiturn path (services/gemini_service.py) and the agent path via a new _build_pro_model_settings() helper threaded into agent.run() whenever the effective model is Pro.
  • Fast lighter: switch "fast" mapping from gemini-2.5-flashgemini-2.5-flash-lite in both routes/learn.py and routes/quiz.py (kept symmetric per ADR 0013).

Why

User-reported: "Smart is taking too long to think." Pro on dynamic thinking can spend 10s+ in the planning phase before streaming any tokens. 2048 tokens is enough for a multi-step pedagogical explanation without burning latency the student can feel — strong reasoning, snappier response. Tunable via _PRO_THINKING_BUDGET in routes/learn.py and the 2048 literal in services/gemini_service.py:114 if we want to dial further.

Fast was already lighter than Smart but still on full Flash. Lite gives meaningfully cheaper + snappier responses for the "I just want a quick answer" use case the tooltip describes.

Behavior matrix

prefmodelthinking
"fast"gemini-2.5-flash-litenone (Lite doesn't think)
"smart"gemini-2.5-procapped at 2048 tokens
(unset)gemini-2.5-pro (agent default)capped at 2048 tokens

Test plan

  • python -m pytest tests/test_learn_routes.py tests/test_quiz_routes.py tests/test_chat_tutor_imports.py -q — 70 pass
  • python -m pytest tests/ -q --ignore=tests/evals — 597 pass; 3 pre-existing live-Supabase failures unchanged (test_skips_self_edges, test_save_to_db, test_full_pipeline — same baseline as PR refactor(quiz): convert generate_quiz to quiz_agent (refactor #2) #71 / refactor(learn): convert chat tutor to chat_tutor_agent (refactor #3) #78)
  • 4 tests rewritten to pin the new mappings: test_fast_returns_lite, test_fast_pref_overrides_agent_model (×2 — learn + quiz), test_legacy_fallback_uses_lite_when_pref_fast
  • Manual: send a Smart-mode tutor message and confirm replies feel snappier without losing reasoning depth
  • Manual: send a Fast-mode tutor message and confirm Lite output still feels useful for quick questions

Notes

  • ADR 0013's symmetry contract between Learn and Quiz routes is preserved — both _PREF_MODEL_NAMES dicts and the legacy quiz fallback all resolve "fast" → Lite identically.
  • No frontend changes needed — model_pref wire format unchanged.
  • Frontend tooltip ("Fast is the default — quicker replies. Switch to Smart for stronger reasoning…") still accurate.

🤖 Generated with Claude Code

Smart was taking too long because Pro ran with dynamic thinking
(`thinking_budget=-1`) — Gemini decided how long to think and could
spend 10s+ in the planning phase before any tokens streamed. Cap it
at 2048 tokens, which keeps multi-step pedagogical reasoning intact
for tutor-length responses while shaving meaningful latency off
every Smart turn. Applied in both the legacy `call_gemini_multiturn`
path (`gemini_service.py:114`) and the agent path via a new
`_build_pro_model_settings()` helper that's threaded into
`agent.run()` whenever the effective model is Pro (explicit "smart"
or no-pref → agent default).
Fast was already lighter than Smart but still on `gemini-2.5-flash`.
Switch the "fast" mapping to `gemini-2.5-flash-lite` (Lite) — same
model the quiz route already uses as its baseline per ADR 0008. Fast
becomes meaningfully cheaper + snappier, with the explicit
"opt-in for speed" semantics the tooltip describes.
Quiz route's `_PREF_MODEL_NAMES` and legacy fallback are updated in
lockstep — ADR 0013 keeps these two routes symmetric so a user
choosing "fast" gets the same model whether they're on Learn or Quiz.
Tests: 4 tests rewritten to assert the new mappings; full backend
suite green except the 3 pre-existing live-Supabase failures (same
baseline as PR #71 / #78).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented May 5, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@Jose-Gael-Cruz-Lopez has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 33 minutes and 48 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 8530bcd9-1334-4d9a-8906-f7f35afb588a

📥 Commits

Reviewing files that changed from the base of the PR and between a8a6243 and abd93a1.

📒 Files selected for processing (8)
  • backend/agents/chat_tutor.py
  • backend/routes/learn.py
  • backend/routes/quiz.py
  • backend/services/gemini_service.py
  • backend/tests/test_gemini_service.py
  • backend/tests/test_learn_routes.py
  • backend/tests/test_quiz_routes.py
  • frontend/src/components/ModelToggle.tsx
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/tutor-thinking-cap

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented May 5, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontendabd93a1Commit Preview URL

Branch Preview URL
May 05 2026, 04:59 AM

Jose-Gael-Cruz-Lopezand others added 2 commits May 5, 2026 00:48
…tings, doc agent sharp edge
- Drop unused MODEL_DEFAULT import from routes/quiz.py (last caller
replaced by MODEL_LITE in this PR; learn.py already pruned theirs).
- Add 3 tests pinning the model_settings contract that previously had
no coverage:
- test_smart_pref_attaches_thinking_cap — asserts budget == 2048
(regression guard against accidentally restoring dynamic thinking)
- test_no_pref_attaches_thinking_cap — confirms no-pref → Pro default
still gets capped, not just explicit "smart"
- test_fast_pref_does_not_attach_thinking_cap — confirms Lite runs
skip thinking_config (it'd be wasted at best)
- Document the route-layer-only enforcement on agents/chat_tutor.py so
any future direct caller of chat_tutor_agent.run(...) knows the cap
isn't on the agent itself.
73 targeted + 581 full suite pass; same 3 pre-existing live-Supabase
failures as the baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add legacy-path thinking-budget tests in test_gemini_service.py:
- test_pro_uses_capped_thinking_budget — pins 2048 on Pro
- test_flash_disables_thinking — pins 0 on Flash + Flash-Lite
Symmetric to the agent-path coverage in test_learn_routes.py so a
refactor can't silently restore Pro to dynamic (-1) on either side.
- Split the chained `budget == _PRO_THINKING_BUDGET == 2048` assertion
into two lines for readability — same regression guarantee.
- Add direct unit test on `_build_pro_model_settings` so the helper's
contract is pinned even if the integration tests above are refactored.
- Update tooltip copy in ModelToggle.tsx — drop "don't mind waiting"
framing now that Smart's thinking is capped (it's snappier than the
old copy implied).
100 targeted + 584 full suite pass; same 3 pre-existing live-Supabase
failures.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jose-Gael-Cruz-Lopez
Jose-Gael-Cruz-Lopez merged commit eedce68 into mainMay 5, 2026
4 checks passed
@AndresL230
AndresL230 deleted the perf/tutor-thinking-cap branch June 27, 2026 04:21
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Jose-Gael-Cruz-Lopez