fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(runtime): reload completed sessions after sandbox boundary decisions - #1609

Merged
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary
Jul 29, 2026
Merged

fix(runtime): reload completed sessions after sandbox boundary decisions#1609
Astro-Han merged 3 commits into
mainfrom
fix/1607-reload-sessions-after-sandbox-boundary

Conversation

@Astro-Han

@Astro-HanAstro-Han commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary

A completed session that contains a sandbox boundary request and decision could not be reopened. AiSdkFlow maps both to control-only stateDelta RuntimeEvents (ai-sdk-flow.ts:303-337) — no content, no recognised action — and projectRuntimeEventsToStoredMessages claimed neither shape, so each produced a hard unsupported_event diagnostic and RuntimeReadModel rejected the entire projection.

The blast radius is wider than the transcript. getSessionView is the sole authority behind getMessages, listTurns, branchFromTurn, branchBeforeTurn, reviseBeforeTurn, requireTurnForAction and requireUserMessageForTurn, so an affected session could not be read, branched, revised, or acted on at all.

The projection now claims a well-formed boundary request or decision and produces no message row for it. The claim matches exactly what AiSdkFlow emits — every field, the system/user identity, and the tool-call reference, with the expansion checked by validateSandboxBoundaryExpansion rather than a local guess. Anything short of that stays an unsupported_event: a partial claim would be the worst of both, paying the cost of rejecting a malformed ledger while still admitting one. Sandbox enforcement, settlement, persistence and the boundary schema are untouched.

Beyond the fix itself, stateDelta is Record<string, unknown> while the projection is a closed whitelist, so nothing forced the two to stay in sync — the actual reason #1581 could introduce this. A projection-coverage contract keyed on BackendSessionEvent['type'] now fails to compile when a variant has no sample, and fails at runtime when the projection does not claim one.

Fixes#1607

Verification

  • packages/runtime full suite: 2767 passed, 0 failed, 9 skipped.
  • Each new test was confirmed to fail for the right reason before its fix, and the read-model change was reverted once to confirm all three layers (projection unit, coverage contract, durable getMessages round trip) turn red together.
  • The coverage table was mutation-tested: a missing entry, an entry without a subject, an empty subject, and a subject borrowed from another variant each fail compilation (TS2741 / TS2322). An earlier revision only caught the first of those.
  • npm run format, npm run lint, npm run build, npm run typecheck across all workspaces: clean.

Root cause

Open-ended write side (stateDelta: Record<string, unknown>), closed whitelist read side, and a hard failure when they disagree. #1581 replaced the whole permission vocabulary across 30+ commits without touching the projection, and nothing caught it.

The failure was also invisible until after the turn ended: an active run is served from the in-flight projection cache (runtime-read-model.ts:98-139) and never re-projects the durable ledger, so approving the boundary, running the tool and rendering the turn all behaved correctly. Only reopening the session hit the ledger path.

Review focus

Two deliberate scope decisions.

No Desktop E2E. The issue suggested one. The sandbox-boundary fixture injects E2eFixtureState directly and never reaches RuntimeEventStore or RuntimeReadModel, and the fake backend cannot emit boundary events, so an E2E along that path would re-test the existing UI takeover assertions (sandbox-boundary-takeover.spec.ts) rather than the regression. The durable SessionManager.getMessages round trip covers it directly, for allow and deny. Making the E2E meaningful would mean building boundary-emission into the fake backend — worth doing on its own terms, not as a smokescreen here.

Non-terminal error content left as is. The coverage contract surfaced it as a third candidate, but AgentRun deliberately keeps non-terminal error RuntimeEvents out of the ledger (agent-run.ts:487), verified against a real SessionManager round trip whose ledger holds only the user text and one terminal failed event. The projection's hard diagnostic there is the matching defensive assertion, so the contract sample reflects the shape a reader actually finds.

Three related gaps found during the analysis are not addressed here and are tracked separately:

  1. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) AiSdkFlow's exhaustiveness guard (ai-sdk-flow.ts:507-521) emits a control-only stateDelta for any unknown SessionEvent, which the projection rejects as hard — so that "safe" fallback escalates one ignorable event into an unreadable session. Whether it should degrade is a policy decision about projection failure handling, not part of this fix. The compile-time coverage contract already catches a new variant well before it can reach a ledger.
  2. (fix: isolate an unclaimed RuntimeEvent instead of rejecting the whole session view #1613) A projection hard-failure has no fallback. backfillMissingRuntimeEvents only fires on an empty ledger, so an intact 10-message session is discarded wholesale over one unclaimed event. Relaxing this trades against its purpose — preventing silently dropped messages — and deserves a separate decision. Same underlying question as (1).
  3. (fix: let a pending sandbox boundary request survive a host restart #1612) A pending boundary request cannot survive a restart. RuntimeKernel.sandboxBoundaryRequestOwners is an in-memory map bound to a backend generation (runtime-kernel.ts:1911), and activePermissionOverlayEvent covers only the permission actions, so the response authority is gone once the host restarts. Recovery then settles every persisted pending request as deny with closureReason: 'host_restarted' (session-manager.ts:915-918) and the turn is marked failed. That is correctly fail-closed, but from the user's side the prompt vanishes and the task is interrupted with no way to resume it.

A completed session containing a sandbox boundary request and decision could
not be reopened. Both are control-only state deltas, the legacy read-model
projection claimed neither shape, and the resulting hard `unsupported_event`
diagnostics made RuntimeReadModel reject the whole projection — which fails
every getSessionView caller, not just the transcript: branching, revising and
any turn-scoped action go with it.
Claim both, and downgrade AiSdkFlow's unmapped-SessionEvent guard to a soft
`unmapped_session_event` diagnostic so one unknown event stays observable
without making a session unreadable.
A projection-coverage contract keyed on BackendSessionEvent['type'] now fails
to compile when a new variant has no sample, and fails at runtime when the
projection does not claim one.
Fixes#1607
Claiming the key alone let any shape ride in under a control-fact name. Only a
well-formed request or decision is a canonical fact; anything else stays an
`unsupported_event`, matching how isPlanProposalStateDelta guards its own claim.
Drop the `unmapped_session_event` downgrade. Turning an unknown event from a
hard failure into a soft diagnostic is a policy decision about projection
failure handling, not part of reloading a session after a boundary decision,
and it would hide a future projection gap. The compile-time coverage contract
already catches a new SessionEvent variant before it can reach a ledger.
Cover deny alongside allow in the durable round trip, and use real boundary
payloads in the projection tests.
…ject
A partial claim was the worst of both: it paid the cost of rejecting a
malformed ledger while still admitting one. A boundary fact now has to match
what AiSdkFlow actually emits — every field, the system/user identity, and the
tool-call reference — with the expansion checked by the authoritative
validator rather than a local guess. Eight table-driven counter-examples cover
the fields, the identity and the reference.
The coverage table only proved its keys existed. `provider_retry: []` compiled
and passed, and the `error` entry had drifted to testing `complete`. Each entry
now holds a `subject` typed to its own key, so neither is expressible, and
companions moved to explicit before/after.
That let `error` hold a real error event again: the contract filters mapped
events through AgentRun's own ledger-admission rule instead of restating it, so
it now says what it means — every event a reader can meet has to project.
Reported by Codex review of #1609.
@Astro-Han
Astro-Han merged commit b981d10 into mainJul 29, 2026
3 checks passed
@Astro-Han
Astro-Han deleted the fix/1607-reload-sessions-after-sandbox-boundary branch July 29, 2026 14:04
Astro-Han added a commit that referenced this pull request Jul 29, 2026
…#1618)
* refactor(runtime): give read-model diagnostics one severity authority
RuntimeReadModel decided which projection diagnostics are fatal by restating
their codes, so the projection declared the diagnostics and its caller declared
what they mean. Move that decision next to the codes as a table keyed by
`RuntimeEventReadModelDiagnosticCode`: a new diagnostic cannot compile without
saying whether it means a user-visible row may be missing. The caller and the
persisted-compat test now ask that authority instead of listing codes.
No behavior change — the table restates today's hard set exactly.
* fix(runtime): isolate an unclaimed control fact from the session view
One RuntimeEvent the projection did not claim made an entire session
unreadable. The catch-all emitted a hard `unsupported_event`, RuntimeReadModel
threw on it, and the whole projection went with it — so getMessages, listTurns,
branching, revising and every turn-scoped action failed over a fact that owns no
chat row. #1607 was one instance; #1609 claimed those two shapes but left the
amplification in place.
Split the catch-all on the RuntimeEvent's own structure: `content` is its
message payload, `actions` its control intent. Every row this projection emits
from an unclaimed shape would have come from content, so a content-bearing
event stays hard — "a message is never silently dropped" is the invariant the
hard failure exists for. A control-only fact has nothing to lose, so it becomes
`unclaimed_control_fact` and degrades the view instead of discarding it. A
projector that tried to build a row and failed still reports its own hard
diagnostic, so this softens nothing that attempted a message.
A future gap is still caught before a user meets it. The projection-coverage
contract now asserts on the unclaimed codes at either severity rather than the
hard one alone, so a new SessionEvent variant with no claim still fails CI, and
AiSdkFlow's exhaustiveness guard is what a variant becomes: a content-free
control fact that lands on the degradable side by construction.
Fixes#1613
* fix(runtime): claim every action field the read-model projection can meet
The soft path rests on a premise that was not machine-checked: an unclaimed
content-free event degrades the view instead of withholding it, which is only
safe while no unclaimed action can owe a row. `content === undefined` does not
prove that on its own — permissionDecision, tokenUsage and the terminal fact all
produce rows, and runtime-event-backfill already writes a content-free event
that becomes a visible `permission_decision`. What actually holds the rule up is
claim coverage, so make coverage the thing that is proven.
The SessionEvent contract only covers events built by
`mapSessionEventToRuntimeEvent`; tool-runtime, terminal-run-commit and the
backfill write RuntimeEvents directly, so a new action field on those paths was
invisible to it. A second contract keyed on `RuntimeEventActions` gives every
field a reachable sample typed to its own key: a new field cannot compile
without one and cannot pass without being claimed.
Writing it found three fields the projection never claimed — `artifactDelta`,
`transferToAgent` and `runtimeProtocol`, the last of which real emitters already
write. All three are control-only, so claim them, and say in the fallback what
the rule actually depends on.
* test(runtime): lock the unclaimed predicate and the caller's hard policy
Two regressions the suite could not see. The unmapped-SessionEvent test compared
the raw code string, so dropping `unclaimed_control_fact` from
`isUnclaimedRuntimeEventDiagnostic` would have quietly narrowed the coverage
contract to `unsupported_event` with every test still green; it now filters
through the predicate itself. And the hard side was only asserted inside the
projector, so a caller that stopped enforcing the policy went unnoticed: append
a content-bearing unclaimed event to a completed run's ledger and getSessionView
must still refuse the view — the counterpart of the soft reproduction beside it.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: reload completed sessions after sandbox boundary decisions

1 participant

@Astro-Han