feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(widget): context-usage meter backed by the agent manager - #58

Merged
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage
Jul 30, 2026
Merged

feat(widget): context-usage meter backed by the agent manager#58
Asaf-prog merged 2 commits into
mainfrom
feat/widget-context-usage

Conversation

@AmitAvital1

@AmitAvital1AmitAvital1 commented Jul 27, 2026

Copy link
Copy Markdown
Collaborator

Stack 1/3 · base: main

Adds a token-budget meter (assistant-ui ContextDisplay shape) to the widget. All computation is on the backend — the agent manager exposes GET /conversations/{id}/usage returning used/max tokens, percent, and a BudgetSeverity (normal/warning/critical) as the single source of truth. The FE only renders; it does no token math. Severity is a StrEnum reused across domain + schema so thresholds live in one place.

The number is the conversation's cumulative token spend (every turn's input + output, summed) measured against context_max_tokens — the budget the service actually enforces with a 429. It is deliberately not a context-window gauge, and the naming says so throughout (review feedback, e4705b0).

Verified: make lint + mypy clean, 472 tests pass, widget typecheck + Playwright e2e green.

🤖 Generated with Claude Code

@Asaf-prog

Copy link
Copy Markdown
Collaborator

PR #58

The implementation is clean, but I think the current naming is misleading.

The backend calculates usage by summing input_tokens + output_tokens across the entire conversation. That represents cumulative token consumption, not the size of the context currently being sent to the model. Since previous messages may be included again in later inputs, the same content can effectively be counted multiple times.

We should either:

  1. Rename this to something like Conversation token budget, or
  2. Calculate the tokens in the current bounded context and compare them against the model’s context-window limit.

I would prefer resolving this before merging, because the UI currently presents the value as Context usage, which gives the user a different meaning from what is actually calculated.

Also, should the critical threshold use >= 85 rather than > 85?

AmitAvital1and others added 2 commits July 28, 2026 20:51
Surface per-conversation token usage against the configured
context_max_tokens budget so users see how full the context is before
they hit the 429 wall.
Backend (source of truth):
- ContextUsage domain value object with a ContextSeverity StrEnum and a
from_totals factory that owns the warning/critical thresholds
- GET /conversations/{id}/usage returns used_tokens, max_tokens, percent
and severity; ConversationService.usage sums stored token counts
- schema reuses the domain ContextSeverity enum (no duplicated literals)
Widget (stateless renderer):
- AgentChatClient.getUsage + useConversation.loadUsage fetch usage on
open and after each turn; no client-side token math or thresholds
- ContextMeter ring (severity-coloured, hover popover) shown only when a
budget is set; hidden and non-breaking against backends without /usage
Also: neutral demo copy and a documented CONTEXT_MAX_TOKENS in the
starter example.
Tests: usage endpoint, severity thresholds, and two widget e2e cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review feedback on #58: the value is a *cumulative* sum of every turn's
input + output tokens, not the size of the context window currently sent
to the model — history is re-sent each turn, so the same message is
counted again every time it is included. Calling it "context usage" told
the user something the number does not mean.
Rename the whole surface to the thing the backend actually enforces (the
per-conversation budget behind ConversationTokenBudgetExceeded / 429):
ContextUsage -> TokenBudgetUsage, ContextSeverity -> BudgetSeverity,
ContextUsageResponse -> TokenBudgetResponse, ContextMeter -> BudgetMeter,
`context-*` CSS -> `budget-*`, and the popover/aria labels now read
"Token budget". The GET .../usage route and CONTEXT_MAX_TOKENS setting
keep their names (pre-existing API), but both are now documented as
cumulative rather than context-window figures.
Also make the critical threshold inclusive (>= 85 like the warning's
>= 65, not > 85) so the two thresholds behave the same, with the boundary
pinned by the test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AmitAvital1
AmitAvital1force-pushed the feat/widget-context-usage branch from 2093eda to e4705b0CompareJuly 28, 2026 18:09
@AmitAvital1
AmitAvital1 changed the base branch from feat/widget-fluent-ui to mainJuly 28, 2026 18:10
@AmitAvital1

Copy link
Copy Markdown
CollaboratorAuthor

Good catch — agreed, and fixed in e4705b0.

You're right about what the number is: it's the cumulative sum of every turn's input_tokens + output_tokens, so history that gets re-sent is counted again each turn. Calling it "context usage" told the user something it doesn't mean.

I went with option 1 (rename) rather than option 2, because the cumulative figure is the one the backend actually enforcesprepare_turn refuses a turn with ConversationTokenBudgetExceeded → 429 once the sum passes context_max_tokens. Measuring the current bounded context against the model's window would show a number nothing acts on, and would leave the real wall invisible. So the meter now names the wall it's warning about:

  • ContextUsageTokenBudgetUsage, ContextSeverityBudgetSeverity, ContextUsageResponseTokenBudgetResponse, ContextMeterBudgetMeter, context-* CSS → budget-*
  • popover reads Token budget, aria-label reads "Token budget N% used"
  • the domain object carries a docstring saying explicitly that this is cumulative, not the context window, and why
  • CONTEXT_MAX_TOKENS in the starter .env.example is now documented the same way (kept the env var and the GET .../usage route names — both pre-existing API)

And yes on the threshold: warning was >= 65 while critical was > 85. Both are inclusive now, with the boundary pinned:

assertTokenBudgetUsage.from_totals(849, 1000).severity=="warning"assertTokenBudgetUsage.from_totals(850, 1000).severity=="critical"

Also rebased onto main (#57 has merged) and set this PR's base to main, so the stack is in order again.

make lint + mypy clean, 472 tests pass, widget typecheck + 17 Playwright e2e green.

@Asaf-prog
Asaf-prog merged commit 160813d into mainJul 30, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui enhancementUser-facing UI improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AmitAvital1@Asaf-prog