fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix: explain a sandbox boundary request closed by a host restart - #1619

Merged
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart
Jul 29, 2026
Merged

fix: explain a sandbox boundary request closed by a host restart#1619
Astro-Han merged 4 commits into
mainfrom
fix/sandbox-boundary-pending-restart

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

When the host restarted while a sandbox boundary request was awaiting the user, the prompt disappeared and the turn ended as a generic failure. Recovery settles every persisted pending request as deny with closureReason: 'host_restarted', which is correctly fail-closed and leaves nothing hanging — but a request the user never saw resolved was closed against them silently, and the turn was lost with no explanation.

This PR delivers the second half of the issue's expectation: the closure is explained. It does not make the request answerable again — the response authority is keyed to a backend generation, and when the process dies so does the turn's execution context, so that needs a resume capability and is tracked separately.

The attribution is owned by the request row. An earlier revision of this PR matched "ids this recovery pass just closed" against the run's RuntimeEvent ledger, and argued for not touching the storage schema. That was wrong, and two reachable orderings broke it:

  • The row is committed before the matching RuntimeEvent is published, and a non-terminal RuntimeEvent append is deliberately fail-open — desktop's default store is JSONL with no canonical marker. Crash in that window and the ledger has no request event, so the attribution silently degraded to a generic restart.
  • If recovery settled the row and then died before the terminal commit, the next startup only queried pending rows and the link was lost for good.

So provenance now lives where the earliest durable fact already is. sandbox_boundary_log gains turn_id and run_id, written in the same INSERT as the request. CreateSandboxBoundaryRequest.turnId is required, so the one real producer cannot forget it. Recovery reads listSandboxBoundaryRestartClosures(sessionId) — rows already settled as denied with host_restarted — and attributes by runId, falling back to turnId only for pre-provenance rows.

This deletes more than it adds: the in-memory Set of closed ids, the extra parameter threaded into recoverAgentRunsFromLedger, the ledger scan and its helper, and the boolean-returning wrapper around the settlement are all gone — −47 lines of production logic. Two properties stopped being checks and became structural: mis-attribution is impossible because the query filters on the settled outcome, so a no-op settlement can never enter the set; and idempotence needed no state bit, because on a later pass the run is already terminal and classifyAgentRunRecovery returns nothing.

Desktop names the failure instead of showing the generic restart message, and offers retry rather than app_restarted's continue, because there is no context left to continue from. The fail-closed deny semantics are unchanged.

Closes#1612. Refs #1607.

Migration

Schema 14 adds the two columns plus a partial index on settled closures. Additive, applies to any v13 database. Legacy rows decode with no provenance and attribute nothing — no backfill, since those requests were settled long ago. There is only one storage path to migrate: FileSessionStore never held these rows and throws 'Sandbox boundary requests require the SQLite metadata control plane'.

Verification

Rebased onto current main (which now carries #1618 and #1610) and re-run there:

SuiteResult
@maka/runtime2804 tests, 2795 pass, 0 fail
@maka/desktop main2948 pass, 0 fail
@maka/storage807 tests, 806 pass, 0 fail
@maka/core1193 pass, 0 fail

The two orderings that broke the old design are now regression tests against real storescreateSessionStore + createAgentRunStore + createRuntimeEventStore over one temp root, driven through three separate store lifetimes (seed, recover, assert), each opening and closing its own handles, so the durable state really does cross a process boundary:

  1. A closure whose RuntimeEvent never reached the ledger still attributes.
  2. A closure re-read across a recovery interrupted before the terminal commit still attributes, and a third pass adds no second turn_state.

Both were verified to be load-bearing: stubbing listSandboxBoundaryRestartClosures to return [] fails them with actual: 'app_restarted', and restoring it makes them pass.

Also locked: an answered request keeps the generic restart class; a reasonless endTurn deny never reads as a restart (plus a store-level test that the query returns only the host_restarted row out of four settlements); a closure belonging to another turn/run is never lent to this one; the boundary overlay keeps its sandboxBoundaryDecision half, not just the request; request provenance is durable and a reuse that changes it is rejected; and a v13 database migrates with legacy rows left without provenance.

npm run typecheck, format:check, lint clean. E2E sandbox-boundary-takeover and session-management pass.

Known gap

No cross-restart E2E. Not because the harness cannot relaunch, but because the fixture's boundary prompt is synthesised in memory: sandboxBoundaryState() injects a SandboxBoundaryRequestEvent into renderer interaction state with no durable row behind it, and seedE2eFixture() re-seeds on every launch, so a second launch would find nothing to recover. A real journey needs durable pending-row seeding, a non-terminal run and ledger for recovery to attribute, first-launch-only seeding, and a stable userDataDir fixture. That is its own piece of work; the real-store integration tests above cover the durable path in the meantime.

Follow-up

"Make a boundary request answerable after a restart" needs its own issue. One constraint worth recording: permission-response-ipc-boundary.test.ts deliberately forbids the renderer from turning an ownerless persisted request row into an actionable prompt. That guard is correct — any future design has to establish a new owner for the persisted request and then relax it, not route around it.

Startup recovery already settled a pending sandbox boundary request
fail-closed (`deny` with `closureReason: 'host_restarted'`), but the fact
stopped at the request row: the interrupted turn recovered as a generic
`app_restarted` failure, so a user who never saw the prompt resolved had
no way to connect the two.
Recovery now keeps the ids it closed and, using the RuntimeEvent ledger
as the only durable map from request id to run and turn, re-attributes
that turn's failure to `sandbox_boundary_closed_by_restart`. An id must
appear in both the closure set and the run's ledger to claim the
attribution, so a stalled child run or an ordinary interrupted turn keeps
its original class.
The active-run overlay also carries sandbox boundary requests and
decisions now, alongside the permission facts it already kept, so a
pending request stays visible in the read-model view instead of vanishing
whenever the messages come from the in-flight projection cache.
The fail-closed deny is unchanged; nothing becomes answerable again.
Refs #1612
…tart
The failed-turn banner reads `sandbox_boundary_closed_by_restart` through
the existing `describeTurnErrorClass` seam, checked before the generic
prefix list so it never degrades into the "waiting for permission" or
"tool failed" catch-alls.
Recovery guidance is `retry`, not the `continue` a plain restart gets:
the request was denied and its backend generation is gone, so there is
nothing to resume into — retrying the turn is the only way to let the
agent ask again.
Refs #1612
…uests
The request row is the earliest durable trace of a boundary prompt — it
commits before the matching RuntimeEvent, whose append is fail-open — but
it recorded nothing about the work it interrupted. Recovery therefore had
to reconstruct that link from the ledger, which a crash in that window
erases.
Schema 14 adds `turn_id` and `run_id` to `sandbox_boundary_log`, plus
`listSandboxBoundaryRestartClosures()`: the settled `denied` +
`host_restarted` rows, indexed and re-readable long after they stop being
pending. That is what a recovery interrupted before its terminal commit
needs on its next attempt.
Because the query filters on the settled outcome, a settlement that was a
no-op — the request had already been approved or plainly denied — can
never be described as a host-restart closure.
Legacy rows read back with no provenance. They are long settled, so
"not attributable" is the honest answer and no backfill is invented.
Refs #1612
The attribution previously rested on the intersection of two volatile
things: a `Set<string>` of ids the current recovery pass happened to
close, and a matching request event in the RuntimeEvent ledger. Two
reachable orderings broke it — a crash between the row commit and the
fail-open event append left no ledger event, and a recovery interrupted
before the terminal commit could never re-derive the set, because the
rows it settled are no longer pending.
ToolRuntime now writes the turn and run onto the request itself, and
recovery reads the settled closures back from the store. The set, the
ledger scan, and the parameter that threaded the set into
`recoverAgentRunsFromLedger` are all gone; a closure claims a run by
`runId`, or by `turnId` for rows predating run identity.
Covered against the real SQLite + JSONL stores across three separate
store lifetimes, including a recovery interrupted after settlement, which
must attribute on the next attempt and stay a no-op on the one after.
Refs #1612
@Astro-Han
Astro-Hanforce-pushed the fix/sandbox-boundary-pending-restart branch from 3173db7 to 67728f6CompareJuly 29, 2026 15:42
@Astro-Han
Astro-Han merged commit e7cd9bf into mainJul 29, 2026
3 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: let a pending sandbox boundary request survive a host restart

1 participant

@Astro-Han