fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(server): refresh Claude context meter after compact - #7249

Open
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3
Open

fix(server): refresh Claude context meter after compact#7249
matsvarn wants to merge 11 commits into
pingdotgg:mainfrom
matsvarn:cursor/claude-compact-meter-fallback-4df3

Conversation

@matsvarn

@matsvarnmatsvarn commented Aug 16, 2026

Copy link
Copy Markdown

Fixes#4650.

What changed

compact_boundary only emitted thread.token-usage.updated when compact_metadata.post_tokens was a finite integer > 0. SDK post_tokens is optional, so a successful /compact often sent no usage event and the composer meter kept the pre-compaction usedTokens.

This keeps that cheap path, then falls back once to the existing queryCurrentContextUsage (getContextUsage) when metadata is missing, zero, or unusable. Compacted thread state is still emitted if neither source yields a snapshot; no usage value is invented.

task_progress / task_notification were writing cumulative total_tokens into usedTokens through Math.max. That cannot decrease, so a correct post-compact reading got ratcheted back up — worse with subagent progress that also carries tool_uses / duration_ms. Those events now keep the last active usedTokens and only advance totalProcessedTokens when a prior reading exists.

apps/server/src/provider/Layers/ClaudeAdapter.ts | 73 +--
apps/server/src/provider/Layers/ClaudeAdapter.test.ts | 510 +++++++++++++++++----
2 files changed, 460 insertions(+), 123 deletions(-)

Why it should exist

After /compact the meter stayed red/full, then could snap down and back up later with no user action. The UI already follows the latest usedTokens; the stale value was coming from the adapter. Teaching the web meter about Claude payloads would still leave lastKnownTokenUsage wrong.

Test

vp test run src/provider/Layers/ClaudeAdapter.test.ts is 76/76.

  • valid post_tokens does not call getContextUsage
  • unusable metadata (undefined / {} / 0 / "18000" / -1) calls it once
  • missing or throwing query emits compacted state and no fabricated usage
  • compact then cumulative task_progress keeps usedTokens: 18000
  • first task_progress no longer invents usedTokens

No UI change, so no before/after images.


Note

Medium Risk
Touches core Claude session lifecycle, token metering, and cleanup on close failures; behavior changes are broad though scoped to the adapter layer with strong test coverage.

Overview
Fixes stale or ratcheting context-meter readings after /compact and during resume by changing how ClaudeAdapter derives and emits thread.token-usage.updated.

Compaction: Valid compact_metadata.post_tokens still drives usage without calling getContextUsage; unusable or oversized metadata falls back once to getContextUsage, and no usage is fabricated if both paths fail. Compacted thread state still emits.

Task progress:task_progress / task_notification no longer map cumulative totals into active usedTokens; they only bump totalProcessedTokens when a prior active reading exists, so post-compact meters are not pushed back up.

Resume & model changes: After a resume handshake, authoritative usage is refreshed (thread-only, no bogus turn.completed). Model/context-window transitions bump modelRevision, gate emissions while setModel runs, rebase cached usage for smaller windows, and retry resume refresh after success or failure.

Session lifecycle:sendTurn is serialized per session (Semaphore + sendTurnLocked); provider controls race a stop signal. Failed query.close() now finishes cleanup (session removed, sends rejected) instead of leaving a “ready” session alive; session.exited can report exitKind: 'error'.

Tests in ClaudeAdapter.test.ts are expanded heavily to lock in these behaviors.

Reviewed by Cursor Bugbot for commit 952478c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Refresh Claude context meter after compact in ClaudeAdapter

  • Fixes token usage tracking after compaction in ClaudeAdapter.ts by reworking normalizeClaudeTaskProgressTokenUsage to only advance totalProcessedTokens monotonically, preventing active token inflation or regression.
  • Introduces ClaudeTokenUsageObservation to tag usage emissions with modelRevision and suppress them during in-flight model transitions.
  • Makes model and context-window changes atomic transitions in sendTurnLocked, using a sendSemaphore to serialize sends and a stopSignal to coordinate shutdown.
  • Validates compact boundary metadata against the known context window in compactBoundaryTokenUsageSnapshot, falling back to an authoritative usage query when metadata is missing, invalid, or exceeds the window.
  • Behavioral Change: task_progress events no longer emit thread.token-usage.updated; stopSessionInternal now uses a stopSignal and an exitPolicy ("always", "on-error", "never") to conditionally emit session.exited.

Macroscope summarized 952478c.

cursoragentand others added 2 commits August 16, 2026 18:17
…ens is missing
Claude compact_boundary events can omit or zero post_tokens, and the adapter
dropped usage updates instead of querying getContextUsage. The composer meter
then kept the pre-compaction usedTokens.
Keep the cheap post_tokens path. Fall back to the existing context-usage query
only when metadata is unusable, and emit no fabricated usage when that query
is absent or fails.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Task progress reports cumulative spend, often from subagents. Writing that
total into usedTokens through Math.max made the context meter unable to
drop after compact and snap back to full. Keep the last active reading and
only advance totalProcessedTokens.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@coderabbitai

coderabbitaiBot commented Aug 16, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 9ab65d93-c034-4596-966c-57715190daf9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 16, 2026
@matsvarn
matsvarn marked this pull request as ready for review August 16, 2026 20:01
@macroscopeapp

macroscopeappBot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The production Claude adapter now changes context metering, model-transition concurrency, resume handling, and terminal session cleanup, including behavior on interruptions and close failures. The extensive tests improve confidence, but the breadth and lifecycle impact make this more than a narrowly scoped bug fix.

You can add or adjust custom eligibility rules. Learn more.

@KanaliserenKanaliseren left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this — separating active context usage from cumulative processed tokens is the right direction, and the compact-boundary fallback addresses the core failure in #4650.

I found three cases worth addressing before merge:

  1. Validate positive post_tokens against the context window.compactBoundaryTokenUsageSnapshot accepts every finite positive value. The shared snapshot builder then clamps oversized values to maxTokens, so an invalid or stale-window value can pin the meter at 100% and prevent the authoritative getContextUsage() fallback. Please treat values above the known context window as unusable and query instead, with a focused test.

  2. Do not let the fallback block the compacted state indefinitely. The compact-boundary handler awaits getContextUsage() before emitting thread.state.changed. Rejection and absence are handled, but a control request that never settles can suppress the compacted event and block later stream processing. The state event should be emitted independently, or the query should have a bounded timeout.

  3. Consider an authoritative refresh when initializing or resuming a Claude session. We observed a real session whose T3 meter showed roughly 3M/full while Claude /context reported 116.6k/1m; running /context immediately corrected T3. This PR refreshes at compact boundaries and turn completion, but it does not appear to guarantee that a resumed session is correct before either event occurs. A one-time getContextUsage() sync on attach/resume would cover that path.

Non-blocking: the five compact-boundary tests repeat most of the same session/stream choreography. A small helper would make the behavior matrix easier to read.

The PR is also currently conflicted with main in ClaudeAdapter.test.ts, so it will need a rebase or conflict resolution regardless.

matsvarnand others added 2 commits August 27, 2026 01:07
# Conflicts:
#	apps/server/src/provider/Layers/ClaudeAdapter.test.ts
Reject oversized compact metadata when Claude has a known context window. Keep the existing query timeout and resume-handshake refresh behavior covered by focused tests.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@matsvarn

Copy link
Copy Markdown
Author

Addressed all three review cases:

  • Oversized post_tokens now triggers the authoritative usage query when the context window is known.
  • Per-turn model changes now update the known context window.
  • The one-second timeout preserves the compacted event and later event order.
  • The resume handshake refreshes usage without emitting turn.completed.

I merged current origin/main normally in d46bd542d. The follow-up fix is a69e2bd4d.

Verified locally:

  • Claude adapter tests: 90 passed.
  • Server typecheck: passed.
  • Targeted lint: passed with one existing warning.
  • Formatting check: passed for both changed files.
  • git diff --check: passed.

I did not verify the full CI suite.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
codexand others added 2 commits August 27, 2026 05:32
Preserve accepted context state, reject stale usage observations, and keep late resume handshakes thread-only.
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

Copy link
Copy Markdown
Author

Update for final head 980e884e. This supersedes the earlier a69e2bd and 90-test note.

  • Resume handshakes now refresh usage without completing a real user turn.
  • Unusable compact metadata uses the bounded one-second authoritative fallback.
  • Oversized post_tokens values are rejected before clamping. Equality with the known limit remains valid.
  • Same-model selections preserve API-learned windows. Real model or window changes clear or rebase stale usage state.
  • Current origin/main was integrated with normal merge commits. No rebase was used.

Local proof:

  • Claude adapter suite: 108 tests passed.
  • Format check passed.
  • Targeted lint passed with one inherited unicorn(no-useless-spread) warning.
  • Server typecheck passed.
  • git diff --check passed.
  • Two independent reviews found no remaining issues.

No live Claude SDK session was tested. CI proof remains pending until CI completes.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit dbe14d8. Configure here.

Comment threadapps/server/src/provider/Layers/ClaudeAdapter.ts
codexand others added 3 commits August 27, 2026 06:39
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
Co-authored-by: Mats Varnskühler <imMxts@users.noreply.github.com>
@github-actionsgithub-actionsBot added size:XL 500-999 changed lines (additions + deletions). and removed size:L 100-499 changed lines (additions + deletions). labels Aug 27, 2026
@matsvarn

matsvarn commented Aug 27, 2026

Copy link
Copy Markdown
Author

@Kanaliseren, this update addresses the three cases in your review.

Final head: 952478cecba233ecb6324201c68b2dd1530f5ac4.

  • Oversized compact post_tokens values are rejected. A bounded authoritative query supplies the fallback.
  • An unresponsive query cannot block the compacted state or later stream events.
  • Resume performs one authoritative refresh. Model revisions reject stale observations.

Additional review fixes preserve API-learned windows and compaction policy during fallbacks. Terminal close cleanup now emits the correct failed-replacement exit. Send and model transitions are serialized.

I merged current main normally through f925d639. The final merge had no conflict. I did not rebase or force-push.

Local proof:

  • ClaudeAdapter tests: 112 passed.
  • New main schema test: 1 passed.
  • Server and Codex schema package typechecks: passed.
  • Format and diff checks: passed.
  • Targeted lint exited 0 with inherited warnings.
  • Diff and worktree checks are clean.
  • Two independent final reviews found no issues.

Remote proof:

  • Macroscope correctness and Effect checks pass.
  • Cursor Bugbot passes with no comments on the final SHA.
  • CodeRabbit skipped its review.
  • Macroscope approvability finished neutral because the review and vouch state is incomplete.

Limits: Tests use a mocked Claude SDK. Local Node is 26.7.0, but the repository requires 24.13.1. No live SDK timing test was run.

Please review and vouch if this is ready.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL500-999 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: context meter ratchets up and never reflects /compact — task_progress writes cumulative tokens into usedTokens via Math.max

4 participants

@matsvarn@Kanaliseren@cursoragent@codex