bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

Description

What happened

A model sent agent_swarm one malformed item (items[2] carried neither
subagent_id nor a legacy profile). That is a recoverable mistake, and the
runtime is built to treat it as one: formatToolArgsViolationText exists so the
refusal names the fields the call does take, and
packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
comes back as a tool result the model can read and correct.

Instead the whole turn died. From the user's side the transcript simply stops
under a red Agent Swarm card:

Tool "agent_swarm" arguments failed validation: [
{
"code": "custom",
"message": "Provide exactly one of subagent_id or legacy profile.",
"path": [ "items", 2 ]
}
] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.

No retry, no further steps, composer idle. The model was never given the chance
to fix its third item.

What the durable state shows

The refused call left nothing in the ledger, and the run never reached a
terminal state:

  • runtime_events for the turn stops at event_seq 102 — the assistant text
    that preceded the call. There is no function_call and no function_response
    for agent_swarm anywhere in the session.
  • tool_journal_events / tool_operations have no row for it either; the last
    entries are the preceding Read operations, cleanly outcome_committed.
  • core_agent_runs still holds "status": "running" for that run (it becomes
    failureClass: "app_restarted" on the next launch), carrying:
"traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

So the append of the synthetic error result was rejected by the ledger, the
rejection threw out of the append, and the turn unwound with the run row stuck
mid-flight.

Root cause

packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
mode — but it only closes it for one of the refusal paths:

// Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
: undefined;

admissionFailure is the only pre-dispatch refusal excluded. Every other one
returns beforeprepareDurableToolAttempt (:1107), which is what calls
commitToolPrepared and actually persists the call on the T1 lane:

refusalsite
arguments failed schema validation:950
loop-gate / repeated failing call:1005, :1026, :1049
permission denial, sandbox boundary:1066, :1077
subagent tool limit:1102

For all of those the tool_start event has already been pushed with an
operationId (id: ${operationId}_call), so AgentRun skips its generic
projection of the call and waits for a commitToolPrepared that never comes.
writeSyntheticToolResult (:728) then finds no durable attempt for the call:

constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

— and writes the result onto the generic lane with a fresh UUID and no
operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
(invocationId, toolCallId), finds no call, and raises orphan_response. The
append throws, and the turn dies.

The UUID in the recorded traceWriteError (rather than a
toolop_<hash>_response id) is the fingerprint of that fallback path.

Scope

Observed via the argument-validation path. The other rows in the table above
share the same shape — pushed tool_start with an operationId, returned before
prepareDurableToolAttempt, synthetic result on the generic lane — so they
should be lethal in the same way; that part is read off the code, not yet
reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

Two secondary observations, separable from the fix:

  1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
    at agent_list. The field list appended by formatToolArgsViolationText is
    the top-level one, which is not where the violation was
    (path: ["items", 2]). Once the turn survives the refusal, this is what
    decides whether the model actually repairs the call.
  2. A single rejected event append killing the turn silently — and leaving the
    run row at status: "running" with no terminal event — is its own weakness,
    independent of what caused the rejection.

How to reproduce

  1. Desktop, a real session (durable runtimeCommitSink active).
  2. Get any tool call whose arguments the tool's own zod schema rejects. A
    cross-field superRefine rule is the reliable way to reach ToolRuntime's
    validation, e.g. an agent_swarm item with neither subagent_id nor
    profile.
  3. The refusal renders under the tool's card and the turn ends there. Afterwards
    the session's runtime_events hold no call/response pair for it, and the run
    row in core_agent_runs carries traceWriteError: … orphan_response … while
    still reading status: "running".

Environment

  • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
  • OS: macOS 26.6 (arm64)
  • Surface: Desktop
  • Node.js: v22.17.0
  • Model: deepseek-v4-pro over an openai-compatible connection

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workinghelp wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
       blocks
      (function() {
      function addCopyButtons() {
      document.querySelectorAll('pre code').forEach(function(codeBlock) {
      if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
      codeBlock.parentElement.setAttribute('data-copy-added', 'true');
      var btn = document.createElement('button');
      btn.textContent = 'Copy';
      btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
      btn.onmouseover = function() { this.style.opacity = '1'; };
      btn.onmouseout = function() { this.style.opacity = '0.7'; };
      btn.onclick = function() {
      navigator.clipboard.writeText(codeBlock.textContent).then(function() {
      btn.textContent = 'Copied!';
      setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
      });
      };
      codeBlock.parentElement.style.position = 'relative';
      codeBlock.parentElement.appendChild(btn);
      });
      }
      addCopyButtons();
      // Re-run on dynamic content
      var observer = new MutationObserver(addCopyButtons);
      observer.observe(document.body, { childList: true, subtree: true });
      })();
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

      Description

      What happened

      A model sent agent_swarm one malformed item (items[2] carried neither
      subagent_id nor a legacy profile). That is a recoverable mistake, and the
      runtime is built to treat it as one: formatToolArgsViolationText exists so the
      refusal names the fields the call does take, and
      packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
      comes back as a tool result the model can read and correct.

      Instead the whole turn died. From the user's side the transcript simply stops
      under a red Agent Swarm card:

      Tool "agent_swarm" arguments failed validation: [
      {
      "code": "custom",
      "message": "Provide exactly one of subagent_id or legacy profile.",
      "path": [ "items", 2 ]
      }
      ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
      

      No retry, no further steps, composer idle. The model was never given the chance
      to fix its third item.

      What the durable state shows

      The refused call left nothing in the ledger, and the run never reached a
      terminal state:

      • runtime_events for the turn stops at event_seq 102 — the assistant text
        that preceded the call. There is no function_call and no function_response
        for agent_swarm anywhere in the session.
      • tool_journal_events / tool_operations have no row for it either; the last
        entries are the preceding Read operations, cleanly outcome_committed.
      • core_agent_runs still holds "status": "running" for that run (it becomes
        failureClass: "app_restarted" on the next launch), carrying:
      "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

      So the append of the synthetic error result was rejected by the ledger, the
      rejection threw out of the append, and the turn unwound with the run row stuck
      mid-flight.

      Root cause

      packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
      mode — but it only closes it for one of the refusal paths:

      // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
      ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
      : undefined;

      admissionFailure is the only pre-dispatch refusal excluded. Every other one
      returns beforeprepareDurableToolAttempt (:1107), which is what calls
      commitToolPrepared and actually persists the call on the T1 lane:

      refusalsite
      arguments failed schema validation:950
      loop-gate / repeated failing call:1005, :1026, :1049
      permission denial, sandbox boundary:1066, :1077
      subagent tool limit:1102

      For all of those the tool_start event has already been pushed with an
      operationId (id: ${operationId}_call), so AgentRun skips its generic
      projection of the call and waits for a commitToolPrepared that never comes.
      writeSyntheticToolResult (:728) then finds no durable attempt for the call:

      constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
      ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

      — and writes the result onto the generic lane with a fresh UUID and no
      operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
      (invocationId, toolCallId), finds no call, and raises orphan_response. The
      append throws, and the turn dies.

      The UUID in the recorded traceWriteError (rather than a
      toolop_<hash>_response id) is the fingerprint of that fallback path.

      Scope

      Observed via the argument-validation path. The other rows in the table above
      share the same shape — pushed tool_start with an operationId, returned before
      prepareDurableToolAttempt, synthetic result on the generic lane — so they
      should be lethal in the same way; that part is read off the code, not yet
      reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
      whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

      Two secondary observations, separable from the fix:

      1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
        at agent_list. The field list appended by formatToolArgsViolationText is
        the top-level one, which is not where the violation was
        (path: ["items", 2]). Once the turn survives the refusal, this is what
        decides whether the model actually repairs the call.
      2. A single rejected event append killing the turn silently — and leaving the
        run row at status: "running" with no terminal event — is its own weakness,
        independent of what caused the rejection.

      How to reproduce

      1. Desktop, a real session (durable runtimeCommitSink active).
      2. Get any tool call whose arguments the tool's own zod schema rejects. A
        cross-field superRefine rule is the reliable way to reach ToolRuntime's
        validation, e.g. an agent_swarm item with neither subagent_id nor
        profile.
      3. The refusal renders under the tool's card and the turn ends there. Afterwards
        the session's runtime_events hold no call/response pair for it, and the run
        row in core_agent_runs carries traceWriteError: … orphan_response … while
        still reading status: "running".

      Environment

      • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
      • OS: macOS 26.6 (arm64)
      • Surface: Desktop
      • Node.js: v22.17.0
      • Model: deepseek-v4-pro over an openai-compatible connection

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        bugSomething isn't workinghelp wantedExtra attention is needed

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

          Description

          What happened

          A model sent agent_swarm one malformed item (items[2] carried neither
          subagent_id nor a legacy profile). That is a recoverable mistake, and the
          runtime is built to treat it as one: formatToolArgsViolationText exists so the
          refusal names the fields the call does take, and
          packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
          comes back as a tool result the model can read and correct.

          Instead the whole turn died. From the user's side the transcript simply stops
          under a red Agent Swarm card:

          Tool "agent_swarm" arguments failed validation: [
          {
          "code": "custom",
          "message": "Provide exactly one of subagent_id or legacy profile.",
          "path": [ "items", 2 ]
          }
          ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
          

          No retry, no further steps, composer idle. The model was never given the chance
          to fix its third item.

          What the durable state shows

          The refused call left nothing in the ledger, and the run never reached a
          terminal state:

          • runtime_events for the turn stops at event_seq 102 — the assistant text
            that preceded the call. There is no function_call and no function_response
            for agent_swarm anywhere in the session.
          • tool_journal_events / tool_operations have no row for it either; the last
            entries are the preceding Read operations, cleanly outcome_committed.
          • core_agent_runs still holds "status": "running" for that run (it becomes
            failureClass: "app_restarted" on the next launch), carrying:
          "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

          So the append of the synthetic error result was rejected by the ledger, the
          rejection threw out of the append, and the turn unwound with the run row stuck
          mid-flight.

          Root cause

          packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
          mode — but it only closes it for one of the refusal paths:

          // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
          ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
          : undefined;

          admissionFailure is the only pre-dispatch refusal excluded. Every other one
          returns beforeprepareDurableToolAttempt (:1107), which is what calls
          commitToolPrepared and actually persists the call on the T1 lane:

          refusalsite
          arguments failed schema validation:950
          loop-gate / repeated failing call:1005, :1026, :1049
          permission denial, sandbox boundary:1066, :1077
          subagent tool limit:1102

          For all of those the tool_start event has already been pushed with an
          operationId (id: ${operationId}_call), so AgentRun skips its generic
          projection of the call and waits for a commitToolPrepared that never comes.
          writeSyntheticToolResult (:728) then finds no durable attempt for the call:

          constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
          ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

          — and writes the result onto the generic lane with a fresh UUID and no
          operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
          (invocationId, toolCallId), finds no call, and raises orphan_response. The
          append throws, and the turn dies.

          The UUID in the recorded traceWriteError (rather than a
          toolop_<hash>_response id) is the fingerprint of that fallback path.

          Scope

          Observed via the argument-validation path. The other rows in the table above
          share the same shape — pushed tool_start with an operationId, returned before
          prepareDurableToolAttempt, synthetic result on the generic lane — so they
          should be lethal in the same way; that part is read off the code, not yet
          reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
          whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

          Two secondary observations, separable from the fix:

          1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
            at agent_list. The field list appended by formatToolArgsViolationText is
            the top-level one, which is not where the violation was
            (path: ["items", 2]). Once the turn survives the refusal, this is what
            decides whether the model actually repairs the call.
          2. A single rejected event append killing the turn silently — and leaving the
            run row at status: "running" with no terminal event — is its own weakness,
            independent of what caused the rejection.

          How to reproduce

          1. Desktop, a real session (durable runtimeCommitSink active).
          2. Get any tool call whose arguments the tool's own zod schema rejects. A
            cross-field superRefine rule is the reliable way to reach ToolRuntime's
            validation, e.g. an agent_swarm item with neither subagent_id nor
            profile.
          3. The refusal renders under the tool's card and the turn ends there. Afterwards
            the session's runtime_events hold no call/response pair for it, and the run
            row in core_agent_runs carries traceWriteError: … orphan_response … while
            still reading status: "running".

          Environment

          • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
          • OS: macOS 26.6 (arm64)
          • Surface: Desktop
          • Node.js: v22.17.0
          • Model: deepseek-v4-pro over an openai-compatible connection

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            bugSomething isn't workinghelp wantedExtra attention is needed

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

              Description

              What happened

              A model sent agent_swarm one malformed item (items[2] carried neither
              subagent_id nor a legacy profile). That is a recoverable mistake, and the
              runtime is built to treat it as one: formatToolArgsViolationText exists so the
              refusal names the fields the call does take, and
              packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
              comes back as a tool result the model can read and correct.

              Instead the whole turn died. From the user's side the transcript simply stops
              under a red Agent Swarm card:

              Tool "agent_swarm" arguments failed validation: [
              {
              "code": "custom",
              "message": "Provide exactly one of subagent_id or legacy profile.",
              "path": [ "items", 2 ]
              }
              ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
              

              No retry, no further steps, composer idle. The model was never given the chance
              to fix its third item.

              What the durable state shows

              The refused call left nothing in the ledger, and the run never reached a
              terminal state:

              • runtime_events for the turn stops at event_seq 102 — the assistant text
                that preceded the call. There is no function_call and no function_response
                for agent_swarm anywhere in the session.
              • tool_journal_events / tool_operations have no row for it either; the last
                entries are the preceding Read operations, cleanly outcome_committed.
              • core_agent_runs still holds "status": "running" for that run (it becomes
                failureClass: "app_restarted" on the next launch), carrying:
              "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

              So the append of the synthetic error result was rejected by the ledger, the
              rejection threw out of the append, and the turn unwound with the run row stuck
              mid-flight.

              Root cause

              packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
              mode — but it only closes it for one of the refusal paths:

              // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
              ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
              : undefined;

              admissionFailure is the only pre-dispatch refusal excluded. Every other one
              returns beforeprepareDurableToolAttempt (:1107), which is what calls
              commitToolPrepared and actually persists the call on the T1 lane:

              refusalsite
              arguments failed schema validation:950
              loop-gate / repeated failing call:1005, :1026, :1049
              permission denial, sandbox boundary:1066, :1077
              subagent tool limit:1102

              For all of those the tool_start event has already been pushed with an
              operationId (id: ${operationId}_call), so AgentRun skips its generic
              projection of the call and waits for a commitToolPrepared that never comes.
              writeSyntheticToolResult (:728) then finds no durable attempt for the call:

              constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
              ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

              — and writes the result onto the generic lane with a fresh UUID and no
              operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
              (invocationId, toolCallId), finds no call, and raises orphan_response. The
              append throws, and the turn dies.

              The UUID in the recorded traceWriteError (rather than a
              toolop_<hash>_response id) is the fingerprint of that fallback path.

              Scope

              Observed via the argument-validation path. The other rows in the table above
              share the same shape — pushed tool_start with an operationId, returned before
              prepareDurableToolAttempt, synthetic result on the generic lane — so they
              should be lethal in the same way; that part is read off the code, not yet
              reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
              whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

              Two secondary observations, separable from the fix:

              1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
                at agent_list. The field list appended by formatToolArgsViolationText is
                the top-level one, which is not where the violation was
                (path: ["items", 2]). Once the turn survives the refusal, this is what
                decides whether the model actually repairs the call.
              2. A single rejected event append killing the turn silently — and leaving the
                run row at status: "running" with no terminal event — is its own weakness,
                independent of what caused the rejection.

              How to reproduce

              1. Desktop, a real session (durable runtimeCommitSink active).
              2. Get any tool call whose arguments the tool's own zod schema rejects. A
                cross-field superRefine rule is the reliable way to reach ToolRuntime's
                validation, e.g. an agent_swarm item with neither subagent_id nor
                profile.
              3. The refusal renders under the tool's card and the turn ends there. Afterwards
                the session's runtime_events hold no call/response pair for it, and the run
                row in core_agent_runs carries traceWriteError: … orphan_response … while
                still reading status: "running".

              Environment

              • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
              • OS: macOS 26.6 (arm64)
              • Surface: Desktop
              • Node.js: v22.17.0
              • Model: deepseek-v4-pro over an openai-compatible connection

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                bugSomething isn't workinghelp wantedExtra attention is needed

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

                  Description

                  What happened

                  A model sent agent_swarm one malformed item (items[2] carried neither
                  subagent_id nor a legacy profile). That is a recoverable mistake, and the
                  runtime is built to treat it as one: formatToolArgsViolationText exists so the
                  refusal names the fields the call does take, and
                  packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
                  comes back as a tool result the model can read and correct.

                  Instead the whole turn died. From the user's side the transcript simply stops
                  under a red Agent Swarm card:

                  Tool "agent_swarm" arguments failed validation: [
                  {
                  "code": "custom",
                  "message": "Provide exactly one of subagent_id or legacy profile.",
                  "path": [ "items", 2 ]
                  }
                  ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
                  

                  No retry, no further steps, composer idle. The model was never given the chance
                  to fix its third item.

                  What the durable state shows

                  The refused call left nothing in the ledger, and the run never reached a
                  terminal state:

                  • runtime_events for the turn stops at event_seq 102 — the assistant text
                    that preceded the call. There is no function_call and no function_response
                    for agent_swarm anywhere in the session.
                  • tool_journal_events / tool_operations have no row for it either; the last
                    entries are the preceding Read operations, cleanly outcome_committed.
                  • core_agent_runs still holds "status": "running" for that run (it becomes
                    failureClass: "app_restarted" on the next launch), carrying:
                  "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

                  So the append of the synthetic error result was rejected by the ledger, the
                  rejection threw out of the append, and the turn unwound with the run row stuck
                  mid-flight.

                  Root cause

                  packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
                  mode — but it only closes it for one of the refusal paths:

                  // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
                  ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
                  : undefined;

                  admissionFailure is the only pre-dispatch refusal excluded. Every other one
                  returns beforeprepareDurableToolAttempt (:1107), which is what calls
                  commitToolPrepared and actually persists the call on the T1 lane:

                  refusalsite
                  arguments failed schema validation:950
                  loop-gate / repeated failing call:1005, :1026, :1049
                  permission denial, sandbox boundary:1066, :1077
                  subagent tool limit:1102

                  For all of those the tool_start event has already been pushed with an
                  operationId (id: ${operationId}_call), so AgentRun skips its generic
                  projection of the call and waits for a commitToolPrepared that never comes.
                  writeSyntheticToolResult (:728) then finds no durable attempt for the call:

                  constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
                  ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

                  — and writes the result onto the generic lane with a fresh UUID and no
                  operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
                  (invocationId, toolCallId), finds no call, and raises orphan_response. The
                  append throws, and the turn dies.

                  The UUID in the recorded traceWriteError (rather than a
                  toolop_<hash>_response id) is the fingerprint of that fallback path.

                  Scope

                  Observed via the argument-validation path. The other rows in the table above
                  share the same shape — pushed tool_start with an operationId, returned before
                  prepareDurableToolAttempt, synthetic result on the generic lane — so they
                  should be lethal in the same way; that part is read off the code, not yet
                  reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
                  whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

                  Two secondary observations, separable from the fix:

                  1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
                    at agent_list. The field list appended by formatToolArgsViolationText is
                    the top-level one, which is not where the violation was
                    (path: ["items", 2]). Once the turn survives the refusal, this is what
                    decides whether the model actually repairs the call.
                  2. A single rejected event append killing the turn silently — and leaving the
                    run row at status: "running" with no terminal event — is its own weakness,
                    independent of what caused the rejection.

                  How to reproduce

                  1. Desktop, a real session (durable runtimeCommitSink active).
                  2. Get any tool call whose arguments the tool's own zod schema rejects. A
                    cross-field superRefine rule is the reliable way to reach ToolRuntime's
                    validation, e.g. an agent_swarm item with neither subagent_id nor
                    profile.
                  3. The refusal renders under the tool's card and the turn ends there. Afterwards
                    the session's runtime_events hold no call/response pair for it, and the run
                    row in core_agent_runs carries traceWriteError: … orphan_response … while
                    still reading status: "running".

                  Environment

                  • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
                  • OS: macOS 26.6 (arm64)
                  • Surface: Desktop
                  • Node.js: v22.17.0
                  • Model: deepseek-v4-pro over an openai-compatible connection

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    bugSomething isn't workinghelp wantedExtra attention is needed

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

                      Description

                      What happened

                      A model sent agent_swarm one malformed item (items[2] carried neither
                      subagent_id nor a legacy profile). That is a recoverable mistake, and the
                      runtime is built to treat it as one: formatToolArgsViolationText exists so the
                      refusal names the fields the call does take, and
                      packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
                      comes back as a tool result the model can read and correct.

                      Instead the whole turn died. From the user's side the transcript simply stops
                      under a red Agent Swarm card:

                      Tool "agent_swarm" arguments failed validation: [
                      {
                      "code": "custom",
                      "message": "Provide exactly one of subagent_id or legacy profile.",
                      "path": [ "items", 2 ]
                      }
                      ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
                      

                      No retry, no further steps, composer idle. The model was never given the chance
                      to fix its third item.

                      What the durable state shows

                      The refused call left nothing in the ledger, and the run never reached a
                      terminal state:

                      • runtime_events for the turn stops at event_seq 102 — the assistant text
                        that preceded the call. There is no function_call and no function_response
                        for agent_swarm anywhere in the session.
                      • tool_journal_events / tool_operations have no row for it either; the last
                        entries are the preceding Read operations, cleanly outcome_committed.
                      • core_agent_runs still holds "status": "running" for that run (it becomes
                        failureClass: "app_restarted" on the next launch), carrying:
                      "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

                      So the append of the synthetic error result was rejected by the ledger, the
                      rejection threw out of the append, and the turn unwound with the run row stuck
                      mid-flight.

                      Root cause

                      packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
                      mode — but it only closes it for one of the refusal paths:

                      // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
                      ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
                      : undefined;

                      admissionFailure is the only pre-dispatch refusal excluded. Every other one
                      returns beforeprepareDurableToolAttempt (:1107), which is what calls
                      commitToolPrepared and actually persists the call on the T1 lane:

                      refusalsite
                      arguments failed schema validation:950
                      loop-gate / repeated failing call:1005, :1026, :1049
                      permission denial, sandbox boundary:1066, :1077
                      subagent tool limit:1102

                      For all of those the tool_start event has already been pushed with an
                      operationId (id: ${operationId}_call), so AgentRun skips its generic
                      projection of the call and waits for a commitToolPrepared that never comes.
                      writeSyntheticToolResult (:728) then finds no durable attempt for the call:

                      constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
                      ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

                      — and writes the result onto the generic lane with a fresh UUID and no
                      operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
                      (invocationId, toolCallId), finds no call, and raises orphan_response. The
                      append throws, and the turn dies.

                      The UUID in the recorded traceWriteError (rather than a
                      toolop_<hash>_response id) is the fingerprint of that fallback path.

                      Scope

                      Observed via the argument-validation path. The other rows in the table above
                      share the same shape — pushed tool_start with an operationId, returned before
                      prepareDurableToolAttempt, synthetic result on the generic lane — so they
                      should be lethal in the same way; that part is read off the code, not yet
                      reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
                      whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

                      Two secondary observations, separable from the fix:

                      1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
                        at agent_list. The field list appended by formatToolArgsViolationText is
                        the top-level one, which is not where the violation was
                        (path: ["items", 2]). Once the turn survives the refusal, this is what
                        decides whether the model actually repairs the call.
                      2. A single rejected event append killing the turn silently — and leaving the
                        run row at status: "running" with no terminal event — is its own weakness,
                        independent of what caused the rejection.

                      How to reproduce

                      1. Desktop, a real session (durable runtimeCommitSink active).
                      2. Get any tool call whose arguments the tool's own zod schema rejects. A
                        cross-field superRefine rule is the reliable way to reach ToolRuntime's
                        validation, e.g. an agent_swarm item with neither subagent_id nor
                        profile.
                      3. The refusal renders under the tool's card and the turn ends there. Afterwards
                        the session's runtime_events hold no call/response pair for it, and the run
                        row in core_agent_runs carries traceWriteError: … orphan_response … while
                        still reading status: "running".

                      Environment

                      • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
                      • OS: macOS 26.6 (arm64)
                      • Surface: Desktop
                      • Node.js: v22.17.0
                      • Model: deepseek-v4-pro over an openai-compatible connection

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        bugSomething isn't workinghelp wantedExtra attention is needed

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

                          Description

                          What happened

                          A model sent agent_swarm one malformed item (items[2] carried neither
                          subagent_id nor a legacy profile). That is a recoverable mistake, and the
                          runtime is built to treat it as one: formatToolArgsViolationText exists so the
                          refusal names the fields the call does take, and
                          packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
                          comes back as a tool result the model can read and correct.

                          Instead the whole turn died. From the user's side the transcript simply stops
                          under a red Agent Swarm card:

                          Tool "agent_swarm" arguments failed validation: [
                          {
                          "code": "custom",
                          "message": "Provide exactly one of subagent_id or legacy profile.",
                          "path": [ "items", 2 ]
                          }
                          ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
                          

                          No retry, no further steps, composer idle. The model was never given the chance
                          to fix its third item.

                          What the durable state shows

                          The refused call left nothing in the ledger, and the run never reached a
                          terminal state:

                          • runtime_events for the turn stops at event_seq 102 — the assistant text
                            that preceded the call. There is no function_call and no function_response
                            for agent_swarm anywhere in the session.
                          • tool_journal_events / tool_operations have no row for it either; the last
                            entries are the preceding Read operations, cleanly outcome_committed.
                          • core_agent_runs still holds "status": "running" for that run (it becomes
                            failureClass: "app_restarted" on the next launch), carrying:
                          "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

                          So the append of the synthetic error result was rejected by the ledger, the
                          rejection threw out of the append, and the turn unwound with the run row stuck
                          mid-flight.

                          Root cause

                          packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
                          mode — but it only closes it for one of the refusal paths:

                          // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
                          ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
                          : undefined;

                          admissionFailure is the only pre-dispatch refusal excluded. Every other one
                          returns beforeprepareDurableToolAttempt (:1107), which is what calls
                          commitToolPrepared and actually persists the call on the T1 lane:

                          refusalsite
                          arguments failed schema validation:950
                          loop-gate / repeated failing call:1005, :1026, :1049
                          permission denial, sandbox boundary:1066, :1077
                          subagent tool limit:1102

                          For all of those the tool_start event has already been pushed with an
                          operationId (id: ${operationId}_call), so AgentRun skips its generic
                          projection of the call and waits for a commitToolPrepared that never comes.
                          writeSyntheticToolResult (:728) then finds no durable attempt for the call:

                          constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
                          ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

                          — and writes the result onto the generic lane with a fresh UUID and no
                          operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
                          (invocationId, toolCallId), finds no call, and raises orphan_response. The
                          append throws, and the turn dies.

                          The UUID in the recorded traceWriteError (rather than a
                          toolop_<hash>_response id) is the fingerprint of that fallback path.

                          Scope

                          Observed via the argument-validation path. The other rows in the table above
                          share the same shape — pushed tool_start with an operationId, returned before
                          prepareDurableToolAttempt, synthetic result on the generic lane — so they
                          should be lethal in the same way; that part is read off the code, not yet
                          reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
                          whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

                          Two secondary observations, separable from the fix:

                          1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
                            at agent_list. The field list appended by formatToolArgsViolationText is
                            the top-level one, which is not where the violation was
                            (path: ["items", 2]). Once the turn survives the refusal, this is what
                            decides whether the model actually repairs the call.
                          2. A single rejected event append killing the turn silently — and leaving the
                            run row at status: "running" with no terminal event — is its own weakness,
                            independent of what caused the rejection.

                          How to reproduce

                          1. Desktop, a real session (durable runtimeCommitSink active).
                          2. Get any tool call whose arguments the tool's own zod schema rejects. A
                            cross-field superRefine rule is the reliable way to reach ToolRuntime's
                            validation, e.g. an agent_swarm item with neither subagent_id nor
                            profile.
                          3. The refusal renders under the tool's card and the turn ends there. Afterwards
                            the session's runtime_events hold no call/response pair for it, and the run
                            row in core_agent_runs carries traceWriteError: … orphan_response … while
                            still reading status: "running".

                          Environment

                          • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
                          • OS: macOS 26.6 (arm64)
                          • Surface: Desktop
                          • Node.js: v22.17.0
                          • Model: deepseek-v4-pro over an openai-compatible connection

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            bugSomething isn't workinghelp wantedExtra attention is needed

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              bug(runtime): every pre-dispatch tool refusal kills the turn — the synthetic result lands as an orphan_response the ledger rejects #2234

                              Description

                              What happened

                              A model sent agent_swarm one malformed item (items[2] carried neither
                              subagent_id nor a legacy profile). That is a recoverable mistake, and the
                              runtime is built to treat it as one: formatToolArgsViolationText exists so the
                              refusal names the fields the call does take, and
                              packages/runtime/src/__tests__/tool-args-violation.test.ts asserts the refusal
                              comes back as a tool result the model can read and correct.

                              Instead the whole turn died. From the user's side the transcript simply stops
                              under a red Agent Swarm card:

                              Tool "agent_swarm" arguments failed validation: [
                              {
                              "code": "custom",
                              "message": "Provide exactly one of subagent_id or legacy profile.",
                              "path": [ "items", 2 ]
                              }
                              ] agent_swarm takes `items`, `prompt_template`, `profile`, `subagent_id`, `resume_run_ids`, `max_concurrency`.
                              

                              No retry, no further steps, composer idle. The model was never given the chance
                              to fix its third item.

                              What the durable state shows

                              The refused call left nothing in the ledger, and the run never reached a
                              terminal state:

                              • runtime_events for the turn stops at event_seq 102 — the assistant text
                                that preceded the call. There is no function_call and no function_response
                                for agent_swarm anywhere in the session.
                              • tool_journal_events / tool_operations have no row for it either; the last
                                entries are the preceding Read operations, cleanly outcome_committed.
                              • core_agent_runs still holds "status": "running" for that run (it becomes
                                failureClass: "app_restarted" on the next launch), carrying:
                              "traceWriteError": "append runtime event: Tool ledger transition rejected: orphan_response at 8de09ddf-1b18-4b0c-964f-8d69be078d69"

                              So the append of the synthetic error result was rejected by the ledger, the
                              rejection threw out of the append, and the turn unwound with the run row stuck
                              mid-flight.

                              Root cause

                              packages/runtime/src/tool-runtime.ts:858-864 already names this exact failure
                              mode — but it only closes it for one of the refusal paths:

                              // Exclusive-step rejection is preflight: it must remain on the generic// call/response lane instead of claiming the T1 dispatch protocol. If the// call carried an operationId here, AgentRun would (correctly) skip its// generic projection assuming commitToolPrepared already persisted it;// the synthetic response would then become an orphan.constoperationId=this.input.runtimeCommitSink&&invocationId&&!admissionFailure
                              ? buildToolOperationId({ invocationId,providerToolCallId: toolUseId})
                              : undefined;

                              admissionFailure is the only pre-dispatch refusal excluded. Every other one
                              returns beforeprepareDurableToolAttempt (:1107), which is what calls
                              commitToolPrepared and actually persists the call on the T1 lane:

                              refusalsite
                              arguments failed schema validation:950
                              loop-gate / repeated failing call:1005, :1026, :1049
                              permission denial, sandbox boundary:1066, :1077
                              subagent tool limit:1102

                              For all of those the tool_start event has already been pushed with an
                              operationId (id: ${operationId}_call), so AgentRun skips its generic
                              projection of the call and waits for a commitToolPrepared that never comes.
                              writeSyntheticToolResult (:728) then finds no durable attempt for the call:

                              constdurableAttempt=this.durableToolAttempts.get(durableAttemptKey(turnId,toolUseId));constdurableOutcome=awaitdurableAttempt?.commitOutcome(content,true);queue.push({type: 'tool_result',id: durableOutcome?.id??this.input.newId(),// ← UUID, not `${operationId}_response`
                              ...(durableOutcome ? {operationId: durableOutcome.operationId} : {}),// ← omitted

                              — and writes the result onto the generic lane with a fresh UUID and no
                              operationId. tool-ledger-scanner.ts:351-372 matches responses to calls by
                              (invocationId, toolCallId), finds no call, and raises orphan_response. The
                              append throws, and the turn dies.

                              The UUID in the recorded traceWriteError (rather than a
                              toolop_<hash>_response id) is the fingerprint of that fallback path.

                              Scope

                              Observed via the argument-validation path. The other rows in the table above
                              share the same shape — pushed tool_start with an operationId, returned before
                              prepareDurableToolAttempt, synthetic result on the generic lane — so they
                              should be lethal in the same way; that part is read off the code, not yet
                              reproduced. If so, the class is "every recoverable pre-dispatch refusal is fatal
                              whenever runtimeCommitSink is active", i.e. in the real Desktop runtime.

                              Two secondary observations, separable from the fix:

                              1. agent_swarm's refusal text says Provide exactly one of subagent_id or legacy profile. without listing the valid profile enum values or pointing
                                at agent_list. The field list appended by formatToolArgsViolationText is
                                the top-level one, which is not where the violation was
                                (path: ["items", 2]). Once the turn survives the refusal, this is what
                                decides whether the model actually repairs the call.
                              2. A single rejected event append killing the turn silently — and leaving the
                                run row at status: "running" with no terminal event — is its own weakness,
                                independent of what caused the rejection.

                              How to reproduce

                              1. Desktop, a real session (durable runtimeCommitSink active).
                              2. Get any tool call whose arguments the tool's own zod schema rejects. A
                                cross-field superRefine rule is the reliable way to reach ToolRuntime's
                                validation, e.g. an agent_swarm item with neither subagent_id nor
                                profile.
                              3. The refusal renders under the tool's card and the turn ends there. Afterwards
                                the session's runtime_events hold no call/response pair for it, and the run
                                row in core_agent_runs carries traceWriteError: … orphan_response … while
                                still reading status: "running".

                              Environment

                              • Maka commit: cef8c44d8 (locally packaged unsigned build, app version 0.1.5)
                              • OS: macOS 26.6 (arm64)
                              • Surface: Desktop
                              • Node.js: v22.17.0
                              • Model: deepseek-v4-pro over an openai-compatible connection

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                bugSomething isn't workinghelp wantedExtra attention is needed

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions