Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
542 changes: 542 additions & 0 deletions docs/t3x/loop/DESIGN.md

Large diffs are not rendered by default.

751 changes: 751 additions & 0 deletions docs/t3x/loop/OPTIONS.md

Large diffs are not rendered by default.

916 changes: 916 additions & 0 deletions docs/t3x/loop/RESEARCH.md

Large diffs are not rendered by default.

152 changes: 152 additions & 0 deletions docs/t3x/loop/SUBAGENTS.md
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,152 @@
# How Claude Code subagents actually work on the wire

Captured 2026-08-02 against `claude` 2.1.220 in `--output-format stream-json` mode — the same mode
T3 Code's `ClaudeAdapter` drives through the Agent SDK. Raw captures are in `captures/`.

This exists because issue #38's root cause was described as "a turn reports `completed` while
background subagents are still streaming activities", and the fix depended on guessing at that
mechanism. It no longer does — the mechanism is reproduced below, event for event.

## The foreground lifecycle

`captures/subagent-foreground.ndjson`. One `general-purpose` subagent running `echo`.

| # | event | `parent_tool_use_id` | notes |
| --- | ------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------- |
| 16 | `assistant` → `tool_use` **`Agent`** | `null` | input carries `subagent_type`, `description`, **`run_in_background`** |
| 17 | `system.task_started` | `null` | `task_id`, **`tool_use_id`**, `subagent_type`, `task_type: "local_agent"`, `prompt` |
| 18 | `user` (the subagent's prompt) | **`toolu_01Hyx…`** | |
| 20 | `system.task_progress` | `null` | `usage {total_tokens, tool_uses, duration_ms}`, `last_tool_name` |
| 21 | `assistant` → `tool_use` `Bash` | **`toolu_01Hyx…`** | the subagent's own tool call |
| 22 | `user` → `tool_result` | **`toolu_01Hyx…`** | the subagent's own tool result |
| 23 | `system.task_updated` | `null` | `patch: {status:"completed", end_time}` |
| 24 | `system.task_notification` | `null` | `status`, `summary`, `output_file`, `usage` |
| 25 | `user` → `tool_result` for the `Agent` call | `null` | the subagent's report re-enters the parent; carries `agentId:` |
| 34 | `result.success` | `null` | turn ends |

### The correlation keys

**`parent_tool_use_id` is the nesting signal, and it is complete.** Every message produced _inside_ a
subagent carries `parent_tool_use_id` = the `Agent` tool_use id that spawned it. Main-agent messages
carry `null`. Nothing else is needed to attribute a message to a subagent — no hooks, no heuristics.

The two id spaces join through `task_started`, which carries **both** `task_id` (`a969007280bdf804b`)
and `tool_use_id` (`toolu_01Hyx…`). So:

```
task_id ←→ task_started.tool_use_id == parent_tool_use_id on nested messages
```

## The backgrounded lifecycle — this is issue #38

`captures/subagent-backgrounded.ndjson`. Same thing with `run_in_background: true`.

```
[29] assistant TOOL_USE Agent run_in_background=True
[30] system.background_tasks_changed ← roster changed
[31] system.task_started a7af6055981f2b9a0
...
[76] result.success <<<<<<<<<< TURN ENDS
[77] result.success <<<<<<<<<< TURN ENDS (second result)
[78] system.background_tasks_changed ← AFTER the turn ended
[79] system.task_updated {task_id:"bshxkpls8", patch:{status:"killed", end_time:…}}
[80] system.task_notification {task_id:"bshxkpls8", status:"stopped"}
```

**Task lifecycle events keep arriving after `result`.** That is exactly the #38 shape, and it is not
a bug in T3 — it is how the protocol works. `run_in_background: true` is the trigger.

In this capture the background task was `killed` because `claude -p` exits when the turn ends. In
T3 the session is long-lived, so the work keeps running instead — which is how thread `3a85bdd3`
produced 1,358 activities across 33 minutes inside a turn the server had already closed.

Note also the **two `result` messages** (`num_turns=2`, then `num_turns=1`). Any logic that assumes
one `result` per turn should be checked against this.

> **SUPERSEDED IN PART, 2026-08-07.** Everything below about the _wire protocol_ is still accurate —
> it is captured, not inferred. But the "what T3 discards" table describes `main` only. Upstream
> landed **#5219 `feat: native subagent & workflow observability`** (`a2ca89aa1`, 2026-08-06, +598
> lines in `ClaudeAdapter.ts`) which already implements every gap this document identifies:
> `parent_tool_use_id` → owning-agent resolution, `task_updated` incl. `is_backgrounded`,
> `background_tasks_changed`, a new `ThreadBackgroundLivenessService`, and — decisively —
> **`backgroundLiveness: "working" | "monitoring" | null` on the thread-shell contract**
> (`packages/contracts/src/orchestration.ts:454`), populated by the same `getThreadShellById` read
> Loop Watch already performs (`ProjectionSnapshotQuery.ts:2336-2342`). CodexAdapter got it too, so
> it is cross-provider.
>
> **Do not build a fork-local open-task roster.** Once the sync lands, guard #15 is one field read:
> `if (shell.backgroundLiveness !== null) return skip("background work in flight")`. Building the
> roster would have been a textbook parallel path — a fork capability duplicating an upstream one,
> silently bypassing its guards. See `docs/t3x/SEAMS.md`.
>
> Note upstream made the same restart tradeoff this design did: the registry is in-memory and empty
> after a restart, on the reasoning that "orphaned background work is not live". So the durable
> `updated_at` timer is still required as the backstop.

## What T3 Code does with all of this today

| wire event | T3 handling | file:line |
| -------------------------- | ---------------------------------------------- | ----------------------- |
| `task_started` | mapped → `task.started` | `ClaudeAdapter.ts:2681` |
| `task_progress` | mapped → `task.progress` + token usage | `:2692` |
| `task_updated` | **dropped** — `case "task_updated": return;` | `:2716` |
| `task_notification` | mapped → `task.completed` | `:2718` |
| `background_tasks_changed` | **swallowed** | `:2597` |
| `parent_tool_use_id` | read **only** to discard subagent token deltas | `:2082` |

So T3 knows a task started and that one finished, but it does **not** maintain an open-task roster,
does **not** know a task was backgrounded, and does **not** attribute any nested message to the
subagent that produced it.

Three fields are being thrown away that answer #38 directly:

1. **`task_updated.patch`** — typed as
`{status?: 'pending'|'running'|'completed'|'failed'|'killed'|'paused', description?, end_time?,
total_paused_ms?, error?, is_backgrounded?}` (`sdk.d.ts:4086-4093`). `is_backgrounded` is the flag
that says "this will outlive the turn". `killed`/`failed` are terminal states `task_notification`
may never report.
2. **`background_tasks_changed`** — the authoritative roster-changed signal.
3. **`parent_tool_use_id`** — free, complete subagent attribution.

## What this means for Loop Watch (#38)

The design currently infers liveness from `projection_threads.updated_at` staleness, because the
premise was that T3 cannot know whether background work is still in flight. **It can.**

Maintaining an open-task set from `task_started` / `task_updated` / `task_notification`, keyed on
`task_id`, gives an exact answer at turn-completion time: _is this run finished, or is it paused with
N subagents still working?_ That is strictly better evidence than a silence timer, and it removes the
worst failure mode in the current design — a nudge fired at a thread that is genuinely mid-flight but
quiet.

It does not replace the idle timer. Two reasons the timer stays:

- The roster is **hot-stream state**. A server restart loses it; `updated_at` is a SQL column that
survives. Keep the timer as the durable backstop and use the roster as a _veto_ on firing.
- A task that is `killed` by a provider crash may never emit `task_notification`, so an open-task set
still needs a TTL. (`task_updated{status:"killed"}` covers the clean case — which is precisely the
event T3 currently drops.)

Recommended revision to #38's guard table: add **"no open tasks in the roster"** as a precondition
for nudging, and surface open-task count in the pill (`Loop paused — 3 subagents working`) instead of
the current binary working/stalled.

The `Stop` hook's `background_tasks` roster (`BackgroundTaskSummary {id, type: shell|subagent|monitor|
workflow, status, description, command?, agent_type?}`) is a second, independent source for the same
answer — but it requires `options.hooks`, which is a fresh edit to a churn-12 file. The wire events
above need **zero** new upstream surface: they are already flowing through a `switch` T3 owns the
arms of.

## Reproducing

```bash
claude -p "Use the Agent tool (subagent_type: general-purpose) to launch exactly ONE subagent. \
Its entire job: run the bash command 'echo hello-from-subagent' and report the output." \
--output-format stream-json --verbose --permission-mode bypassPermissions \
--model claude-haiku-4-5-20251001 < /dev/null > sub.ndjson
```

Add `with run_in_background: true` and "do not wait for it" to the prompt for the #38 shape.

Say **Agent tool**, not "Task tool" — the first attempt at this capture said "Task" and the model
reached for `TaskCreate` (the todo-list tool) instead of spawning anything.
81 changes: 81 additions & 0 deletions docs/t3x/loop/captures/subagent-backgrounded.ndjson

Large diffs are not rendered by default.

35 changes: 35 additions & 0 deletions docs/t3x/loop/captures/subagent-foreground.ndjson

Large diffs are not rendered by default.

Loading
Loading