fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5) - #13

Merged
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns
Jul 25, 2026
Merged

fix(server): reconcile crash-frozen running turns to interrupted on boot (#5)#13
radroid merged 1 commit into
mainfrom
t3x/crash-recovery-reconcile-running-turns

Conversation

@radroid

Copy link
Copy Markdown
Owner

Summary

Fixes#5. When the desktop backend exits ungracefully (SIGKILL / crash / OOM — no SIGTERM), in-flight agent turns are left frozen at projection_turns.state = "running" with projection_thread_sessions.active_turn_id still set. Nothing reconciles them on next launch, so on a later resume the provider reports the cut-off turn as cancelled/aborted, and auto-resume never recovers them. A clean quit does the right thing via an Effect finalizer (ProviderService.runStopAlladapter.stopAll() → a session-set that settles running turns to interrupted); only ungraceful exit was broken.

Approach

At fresh boot no turn is actually live — adapter session maps start empty and runtime rows carry no pid/epoch — so every leftover in-flight turn in the read model is by definition a crash remnant. The fix adds a one-shot boot reconciler that mirrors the graceful path exactly:

  • New apps/server/src/orchestration/Layers/CrashRecoveryReconciler.tsreconcileInterruptedTurnsOnBoot reads the read model (ProjectionSnapshotQuery.getSnapshot()), finds threads with a leftover in-flight turn (latestTurn.state === "running", or session status ∈ {running, starting}, or non-null activeTurnId), and for each dispatches a thread.session.set command with status: "stopped", activeTurnId: null (mirroring what ProviderRuntimeIngestion dispatches on session.exited). This settles the running turn to the canonical resumable interrupted state and clears active_turn_id.
  • Wired into serverRuntimeStartup.ts as a startup phase before reactors.start and before command-ready is signalled, so no reactor or user command observes the stale state, and the (now-unblocked) ProviderSessionReaper can idle-sweep the runtime row.

Key correctness property: reconciliation goes through the event log (a dispatched command), never a direct projection write — projections are a pure function of the event log and a direct write would be reverted on rebuild. It is fully error-swallowed (never error channel) so a bad thread or persistence hiccup can't block boot.

Scope (deliberate)

Reconciling to a resumable interrupted state satisfies the issue's primary acceptance ("or at minimum settled to a resumable interrupted state"). Auto-resume re-arming (auto-continuing the turn) needs a new crash marker + a boot-time producer to distinguish crash-interrupted from user/stop-interrupted turns; that's a documented, separate follow-up (noted in code comments) rather than bundled here.

Tests (TDD — red first)

  • CrashRecoveryReconciler.test.ts (new, integration over the real engine + projection pipeline + in-memory sqlite): precondition (turn running, session running, active_turn_id set) → reconcile → turn interrupted with completed_at, session stopped, active_turn_id null, reconciledCount === 1; idempotency (second run is a no-op); multi-thread selectivity (only crashed threads change; already-settled untouched); event-log backing (a new thread.session-set event is appended, surviving a rebuild).
  • serverRuntimeStartup.test.ts (extended): the reconcile phase runs before reactors.start / command-ready, and a dispatch failure is swallowed (startup still completes).
  • t3x/autoResume/guards.test.ts (extended): a settled interrupted + stopped thread is notthreadIsProgressing (documents resumability).

Verification

  • New/extended trio → 33 pass; full src/orchestration/ + src/t3x/autoResume/255 pass.
  • tsgo --noEmit → exit 0, no errors.
  • vp fmt --check on changed files → clean.

Note: src/t3x/autoResume/Reactor.test.ts has a pre-existing timing-flaky case (unchanged by this PR — verified identical to main); unrelated to this change.

🤖 Generated with Claude Code

… boot
Fixes#5.
When the backend exits ungracefully (SIGKILL/crash/OOM, no SIGTERM), the
Effect finalizer a clean quit runs (ProviderService.runStopAll ->
adapter.stopAll() -> a session-set that settles running turns to
"interrupted") never fires. In-flight turns are left frozen at
projection_turns.state="running" with active_turn_id still set; on a later
resume the provider reports the cut-off turn as cancelled/aborted, and
auto-resume never recovers them.
At fresh boot no turn is actually live (adapter session maps start empty;
runtime rows carry no pid/epoch), so every leftover in-flight turn is a crash
remnant. Add a one-shot boot reconciler that mirrors the graceful path: for
each such thread it dispatches a stopped `thread.session.set` command (through
the event log, not a direct projection write), settling the running turn to
the resumable "interrupted" state and clearing active_turn_id. Wired as a
startup phase before reactors start and before commands are accepted.
Scope: reconciling to a resumable interrupted state satisfies the issue's
primary acceptance. Auto-resume RE-ARMING (a crash marker + boot producer to
auto-continue the turn) is left as a documented follow-up.
Adds CrashRecoveryReconciler.test.ts (precondition, reconcile, idempotency,
multi-thread selectivity, event-log backing) and extends serverRuntimeStartup
and autoResume/guards tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c4e9af4f-873f-4aac-99a2-0abf74fde18b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch t3x/crash-recovery-reconcile-running-turns

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@radroid
radroid merged commit 488776d into mainJul 25, 2026
1 check passed
@radroid
radroid deleted the t3x/crash-recovery-reconcile-running-turns branch July 30, 2026 14:23
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ungraceful backend exit leaves in-flight turns frozen as running → surface as cancelled; auto-resume doesn't recover them

1 participant

@radroid