fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(usage): canonical LiteLLM keys must win rate-table collisions - #8548

Closed
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision
Closed

fix(usage): canonical LiteLLM keys must win rate-table collisions#8548
derektrimm wants to merge 1 commit into
pingdotgg:mainfrom
derektrimm:fix/usage-rate-table-collision

Conversation

@derektrimm

@derektrimmderektrimm commented Aug 28, 2026

Copy link
Copy Markdown

What Changed

parseRateTable now gives a canonical (slash-free) LiteLLM key precedence over prefix-stripped keys that normalize to the same name. A prefix-stripped entry only fills a slot no canonical entry claims. Added usagePricing.test.ts (the file did not exist) covering both collision orders, differing base rates, prefixed-only models, the missing-cache-fields fallback, dotted Bedrock-style keys, and end-to-end cost + savings arithmetic in the exact #8534 shape.

Why

The rate table is built last-entry-wins over normalized names. In today's LiteLLM document the canonical claude-fable-5 / claude-opus-5 / claude-opus-4-8 / claude-sonnet-5 entries are shadowed by deepinfra/anthropic/* reseller entries that carry no cache_read_input_token_cost, so the ?? input fallback priced every cache read at full input rate and cache savings computed as zero. gpt-5 / gpt-5-mini / gpt-5-nano lose the same way to replicate/openai/*.

The reporter's screenshot in #8534 confirms this to the cent: fable-5 averages $10.14/M (full input rate) on a bucket that is overwhelmingly cache reads, the six model costs sum exactly to the $62,540.18 headline, and the $9.45 cache savings is producible only by the two models whose canonical keys happen to serialize last. Full arithmetic in the issue comment.

Deliberately unchanged: the ?? input fallback (correct for models that genuinely publish no cache pricing) and stripped-vs-stripped ordering for names with no canonical entry.

Note: beyond the cache column, this also corrects base rates for models whose canonical entries were shadowed by gateway variants with different prices (e.g. gemini-2.5-pro was priced at 2x from a vercel_ai_gateway/* entry). Every retarget moves toward the canonical price.

Fixes#8534

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (server-side pricing only; no UI change)
  • I included a video for animation/interaction changes (n/a)

Note

Medium Risk
Changes reported usage cost and cache-savings math for models affected by LiteLLM key collisions; no auth or payment flow changes, but billing visibility shifts toward canonical rates.

Overview
Fixes #8534, where reseller/gateway LiteLLM keys (e.g. deepinfra/anthropic/claude-fable-5) that normalize to the same model name were overwriting canonical slash-free entries, stripping cache discount rates and sometimes doubling base input/output prices.

parseRateTable now treats keys without / as canonical: they always claim their normalized slot, and prefixed keys are skipped once a canonical name exists; prefixed-only models still populate the table when no canonical key is present.

Adds usagePricing.test.ts with collision ordering, gateway vs canonical base rates, cache fallbacks, dot-prefixed Bedrock-style keys, and end-to-end priceUsage / cacheSavingsUsd checks for the fable-5 scenario.

Reviewed by Cursor Bugbot for commit 9b49df8. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix parseRateTable to let canonical LiteLLM keys win rate-table collisions

  • parseRateTable now tracks canonical (slash-free) model names in a Set and skips prefixed entries (e.g. deepinfra/anthropic/claude-fable-5) when a canonical entry for the same normalized name already exists.
  • Cache read/creation costs default to the input cost when the rate entry omits them.
  • Adds tests in usagePricing.test.ts covering canonical-vs-prefixed collisions, cache-rate defaults, dot-prefixed names, and savings calculations.
  • Behavioral Change: prefixed entries that previously overwrote canonical rates for the same normalized name are now ignored; consumers relying on prefixed-override behavior will see the canonical rates instead.

Macroscope summarized 9b49df8.

@coderabbitai

coderabbitaiBot commented Aug 28, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d48f4f90-9193-493a-b2c8-4b709a4b3450

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 28, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production change alters model-derived usage-dollar and cache-savings calculations across the usage aggregation path. Although the implementation is small and regression-tested, pricing and metering behavior warrants human review.

You can add or adjust custom eligibility rules. Learn more.

parseRateTable strips provider prefixes for lookup, so several LiteLLM
keys collapse onto one normalized name and the last entry in the
document silently wins. For claude-fable-5, claude-opus-5,
claude-opus-4-8 and claude-sonnet-5 that winner is a cache-less
deepinfra/anthropic/* reseller entry, so the ?? input fallback priced
every cache read at full input rate and cache savings computed as zero.
gpt-5, gpt-5-mini and gpt-5-nano lose the same way to replicate/openai/*,
and some base rates were shadowed too (gemini-2.5-pro at 2x from a
vercel_ai_gateway entry).
A slash-free key is the canonical entry for the name transcripts
actually record, so it now always wins its slot; a prefix-stripped key
only fills a slot no canonical entry claims. The ?? input fallback is
unchanged: it is correct for models that genuinely publish no cache
pricing.
Fixespingdotgg#8534
@derektrimm
derektrimmforce-pushed the fix/usage-rate-table-collision branch from 651eb57 to 9b49df8CompareAugust 28, 2026 23:39
@derektrimm

Copy link
Copy Markdown
Author

Closing - #8806 landed a more complete version of this fix (preserving qualified keys rather than collapsing them) and users on #8534 confirm 0.0.37 resolves it. Glad it's fixed.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Usage View shows total usage cost without splitting off the cache savings

1 participant

@derektrimm