fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(usage): price new models without waiting a day for the rate table - #9202

Merged
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh
Sep 3, 2026
Merged

fix(usage): price new models without waiting a day for the rate table#9202
t3dotgg merged 2 commits into
mainfrom
fix/usage-rates-refresh

Conversation

@t3dotgg

@t3dotggt3dotgg commented Sep 2, 2026

Copy link
Copy Markdown
Member

Claude Fable 5.1 usage showed $0.00 on the Usage page. The server fetches LiteLLM's rate table once every 24 hours, so a model that LiteLLM lists after the last fetch stays unpriced until the TTL runs out. Pressing refresh rescanned transcripts but kept the stale table. Claude Code's 1M context variant, written as `claude-fable-5-1[1m]`, never matched the table at all.

Two changes:

  • The refresh button and pull-to-refresh now send `refreshRates` to every connected environment. Each server refetches the table inside the TTL, with a one minute floor so repeated presses do not hammer the source. Range changes and initial loads keep the daily cadence.
  • Rate lookup strips a bracketed suffix such as `[1m]`, so the 1M tier prices at the base rate. That matches the existing rule of pricing at the base tier.

Covers web, desktop, and mobile. Tests cover the suffix lookup and the forced refetch inside the TTL. User docs updated.

Made with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Note

Medium Risk
Changes LiteLLM fetch timing and model rate matching for usage cost; bounded by TTL floor and single-flight locking, but refresh can increase outbound fetches when users hammer refresh.

Overview
Fixes $0.00 usage for newly listed models (e.g. Claude Fable 5.1) and for Claude Code’s [1m] context-tier model strings.

Server: Adds serverRefreshUsageRates RPC and UsageService.refreshRates, which refetches the LiteLLM table inside the normal 24h TTL (with a 1-minute floor and a semaphore so bursts share one HTTP fetch). Normal readSummary scans still use the daily cadence via ensureRates(false).

Clients (web + mobile): Usage refresh / pull-to-refresh now calls refreshUsageRates per environment first (failures ignored), then **refresh**es the usage summary so transcripts rescans pick up updated pricing.

Pricing lookup:stripVariantSuffix strips bracket suffixes (e.g. claude-fable-5-1[1m]) so rates match the base model name.

Tests and user usage docs updated.

Reviewed by Cursor Bugbot for commit d378d2e. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add on-demand usage rate refresh so new models price without waiting a day

  • Adds the server.refreshUsageRates WebSocket RPC and UsageService.refreshRates effect, which force a rate-table load bypassing the 24h TTL but respecting a new RATES_REFRESH_FLOOR_MS one-minute floor.
  • Mobile and web useUsage hooks now trigger the rate-refresh RPC for each connected environment before rescanning usage summaries, and still rescan even when the pricing request fails.
  • Concurrent rate loads are serialized through a single-permit semaphore in UsageService.make; client-side refresh commands use single-flight per environment.
  • lookupRate now strips bracketed variant suffixes (e.g. context-tier) so a variant model resolves to its base model's rate.
  • Risk: RATES_REFRESH_FLOOR_MS in UsageService.ts caps forced refreshes at once per minute; if the floor is too short, repeated refreshes could increase rate-table fetch load on the pricing source.

Macroscope summarized d378d2e.

The usage page fetches LiteLLM's rate table once every 24 hours. A model
released after the last fetch shows $0.00 until the TTL runs out, and the
refresh button did not help. Claude Code's 1M context variant
(claude-fable-5-1[1m]) never matched the table at all.
Refresh now sends refreshRates to every environment, which refetches the
table inside the TTL with a one minute floor. Rate lookup strips a
bracketed variant suffix so the 1M tier prices at the base rate.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@github-actions

github-actionsBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ProviderMetricMain baselineThis PRImpactPR ceiling
CodexTotal thread wire13.4 KiB13.3 KiB−167 B (−1.2%)15.1 KiB
CodexThread snapshot wire6.9 KiB6.9 KiB−5 B (−0.1%)7.3 KiB
CodexLive turn WebSocket wire6.6 KiB6.4 KiB−162 B (−2.4%)7.8 KiB
CodexLive turn WebSocket decoded57.0 KiB55.6 KiB−1.4 KiB (−2.5%)66.4 KiB
CodexLive turn messages10100 (0.0%)21
ClaudeTotal thread wire13.3 KiB13.2 KiB−84 B (−0.6%)15.1 KiB
ClaudeThread snapshot wire6.9 KiB6.9 KiB−10 B (−0.1%)7.3 KiB
ClaudeLive turn WebSocket wire6.4 KiB6.3 KiB−74 B (−1.1%)7.8 KiB
ClaudeLive turn WebSocket decoded56.4 KiB56.3 KiB−88 B (−0.2%)66.4 KiB
ClaudeLive turn messages108−2 (−20.0%)21

Baseline: 70cd258 · PR result: d378d2e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.4 KiB
  • Claude decoded thread snapshot: 110.1 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Comment threadapps/mobile/src/features/usage/UsageRouteScreen.tsx Outdated
Comment threadapps/server/src/usage/UsageService.ts
Comment threadapps/server/src/usage/UsageService.ts Outdated
Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4a76454. Configure here.

Comment threadapps/web/src/components/usage/UsagePage.tsx Outdated
@macroscopeapp

macroscopeappBot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR changes the usage refresh behavior, outbound rate-table fetching, and model pricing results across web, mobile, and server paths. It also modifies the server authorization mapping, which is a sensitive auth-directory change requiring human review.

You can add or adjust custom eligibility rules. Learn more.

…e page
Keying the summary query on refreshRates created new atoms with no value,
so the first refresh dropped the page into its skeleton. The flag also
stuck, so every revalidation forced a fetch.
Refresh now calls server.refreshUsageRates per environment and then
revalidates the existing summary query in place. Forced fetches take a
lock so a burst shares one request, and the one minute floor applies to
a table loaded from disk too.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 2, 2026
@t3dotgg
t3dotgg merged commit 2b745ef into mainSep 3, 2026
27 checks passed
@t3dotgg
t3dotgg deleted the fix/usage-rates-refresh branch September 3, 2026 09:46
github-actionsBot added a commit to omarcresp/t3code-flake that referenced this pull request Sep 3, 2026
## What's Changed
* fix(mobile): show an error instead of an endless preview spinner by @t3dotgg in pingdotgg/t3code#9123
* fix(usage): price new models without waiting a day for the rate table by @t3dotgg in pingdotgg/t3code#9202
* fix(web): unlock the composer when preview capture fails by @t3dotgg in pingdotgg/t3code#9127
* fix(antigravity): refresh the model manifest so older Gemini models fold as legacy by @t3dotgg in pingdotgg/t3code#9397
* perf(ci): reuse dependency checks in release builds by @t3dotgg in pingdotgg/t3code#9399
**Full Changelog**: pingdotgg/t3code@v0.0.39-nightly.20260903.1268...v0.0.39-nightly.20260903.1270
Upstream release: https://github.com/pingdotgg/t3code/releases/tag/v0.0.39-nightly.20260903.1270
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L100-499 changed lines (additions + deletions).vouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@t3dotgg