fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(server): stop actually stops a runaway Claude turn - #3

Merged
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop
Aug 17, 2026
Merged

fix(server): stop actually stops a runaway Claude turn#3
xiaogwu merged 1 commit into
integrationfrom
server/stop-turn-hard-stop

Conversation

@xiaogwu

Copy link
Copy Markdown
Owner

The problem

Clicking Stop during a Claude turn could freeze every thread in the app for close to two minutes, and then not stop the agent at all.

Two defects combine:

  1. ClaudeAdapter.interruptTurn awaited context.query.interrupt() with no time bound. The Claude Agent SDK's interrupt is a control request that can block until the in-flight tool call returns. Measured twice on 2026-08-17 against claudeAgent: 114s and 111.5s.
  2. ProviderCommandReactor drains every thread command through a single DrainableWorker lane. One hung provider call therefore blocks all threads, on every client. Later Stop clicks were persisted to orchestration_events and never processed.

An acknowledged interrupt is also not proof the turn ended. The event log shows the agent kept working straight through the stall: a burst of identically-timestamped tool calls flushed the instant the queue cleared.

The fix

Stop now escalates instead of trusting the provider.

interruptTurn arms a detached watchdog, then asks the SDK to interrupt with a 5 second bound on the ack. If the same turn is still live 5 seconds later, the session is torn down through the existing stopSessionInternal. That path routes through completeTurn, which updates the resume cursor, so a forced teardown still leaves the thread resumable and the UI honest about the turn being interrupted.

Turn identity (context.turnState?.turnId compared against the id captured at interrupt time) is what distinguishes "same turn still live" from "turn ended" and "a new turn started". No new signal or Deferred was needed.

The watchdog runs on a make-level Effect.runForkWith capture, so under test it inherits TestClock and the escalation is deterministic rather than timing-dependent.

Deliberate non-change

The reactor's queue stays serial. Forking the interrupt off the shared lane would let a following turn-start-requested race ahead of a session teardown and prompt into a dying session. Bounding the adapter call shrinks the global block from ~112s to ~5s, which is the win without the race.

Scope

Only the Claude adapter changes. It is the only provider with observed hangs. Codex, Cursor, Grok, Gemini and OpenCode are not audited yet.

No client change is needed. Web, desktop and mobile all send the same thread.turn.interrupt request through packages/client-runtime/src/operations/commands.ts, so all three surfaces are fixed by the server change alone.

Verification

  • vp test run apps/server/src/provider/Layers/ClaudeAdapter.test.ts — 72 passed, including two new tests.
  • Negative check: with the fix reverted, the new never-acknowledged-interrupt test hangs to the 60s harness timeout. That is the bug reproducing.
  • tsgo --noEmit on apps/server — 0 errors.
  • vp lint on both files — clean (one no-useless-spread warning, pre-existing on main).
  • Not verified at runtime in a real client.

Upstream

Upstream knows about this bug and has not shipped a fix. Open issues #4524 and #4713 describe it. Open PR #5891 fixes it by always closing the session and never calling interrupt() at all, which discards the session even on a Stop that would have worked. This fork keeps the graceful path and escalates only when the agent ignores it. If pingdotgg#5891 merges, expect a conflict in these two files.

Written by Claude Opus 5 in T3 Code.

Clicking Stop during a Claude turn could freeze every thread in the app
for close to two minutes, and then not stop the agent at all.
Two defects combined:
1. `ClaudeAdapter.interruptTurn` awaited `context.query.interrupt()` with
no time bound. The SDK's interrupt is a control request that can block
until the in-flight tool call returns. Measured twice on 2026-08-17:
114s and 111.5s.
2. `ProviderCommandReactor` drains every thread command through a single
`DrainableWorker` lane, so that one hung call blocked all threads.
Further Stop clicks were persisted and never processed.
An acknowledged interrupt is also not proof the turn ended. The event log
showed the agent kept working straight through the stall: a burst of
identically-timestamped tool calls flushed the moment the queue cleared.
So Stop now escalates. It arms a detached watchdog, asks the SDK to
interrupt with a 5 second bound on the ack, and if the same turn is still
live 5 seconds later it tears the session down through the existing
`stopSessionInternal`. That path routes through `completeTurn`, which
updates the resume cursor, so a forced teardown still leaves the thread
resumable.
The reactor's queue stays serial on purpose. Forking the interrupt would
let a following `turn-start-requested` race ahead of a session teardown
and prompt into a dying session. Bounding the adapter call shrinks the
global block from ~112s to ~5s, which is the win without the race.
Verified: 72 tests pass in ClaudeAdapter.test.ts. The new
never-acknowledged-interrupt test hangs to the 60s harness timeout
without the fix.
Only the Claude adapter is changed; it is the only provider with
observed hangs. Codex, Cursor, Grok and OpenCode are not audited yet.
Written by Claude Opus 5 in T3 Code.
@github-actionsgithub-actionsBot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Aug 17, 2026
@xiaogwu
xiaogwu marked this pull request as ready for review August 17, 2026 02:53
@xiaogwu
xiaogwu merged commit 5d50f74 into integrationAug 17, 2026
11 of 15 checks passed
xiaogwu pushed a commit that referenced this pull request Aug 17, 2026
….1116
origin/integration carried the two fork PRs (#3 Stop hard-stop, #4 palette
search scroll); local integration carried the nightly 1116 merge. Neither
had both. No conflicts: the sides touch different files.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:Mvouch:trustedPR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@xiaogwu