Uh oh!
There was an error while loading. Please reload this page.
feat(web): surface orchestration agents in the timeline - #4664
feat(web): surface orchestration agents in the timeline#4664shivamhwp wants to merge 333 commits into
Conversation
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Team Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Warning Your free Security trial is over. An organization admin can upgrade to Advanced for continuous pull request security review or dismiss this notice. Comment |
940e8fb to
c3a50f8Compared2ef6c7 to
adb1e9bComparec3a50f8 to
8c6a796Compareadb1e9b to
39912b0CompareUh oh!
There was an error while loading. Please reload this page.
8c6a796 to
103b937Compare39912b0 to
6bc3a8dCompare103b937 to
31adb36Compare6bc3a8d to
05fe44cCompare31adb36 to
b270c47Compare05fe44c to
7458298Compareb270c47 to
daf1a88Compare7458298 to
4daeac4Comparedaf1a88 to
1dce464Compare4daeac4 to
095cab7Compare1dce464 to
4c169edCompare095cab7 to
1d37a6fCompare4c169ed to
70d5111Compare1d37a6f to
6cf67a2Compare70d5111 to
c78d0ffCompare6cf67a2 to
5e3c535Comparec78d0ff to
a997806Compare5e3c535 to
2430736CompareUh oh!
There was an error while loading. Please reload this page.
Keep typed plan and todo records distinct from generic tool events, activate captured plans, and supersede older planning state within the owning thread. Ignore nested todo snapshots for the parent and retain identity across duplicate SDK messages. Finding: R05 follow-up Model: GPT-5.6 Sol via Codex
R04 follow-up Convert client multi-select answer arrays to the comma-separated string shape required by the pinned Claude SDK while preserving single-select strings. Implemented by GPT-5.6 Sol via Codex.
Stale cached user and status events could admit and complete a newly submitted OpenCode prompt. Generate the native message ID before submission and only advance admission when that exact message is observed. Finding: R14 Implemented by GPT-5.6 Sol with Codex.
R12 follow-up Insert persistent feedback blocks by their timestamp within the canonical timeline while preserving projected row order. Keep real optimistic sends appended and suppress duplicate local messages already committed by the server. Implemented by GPT-5.6 Sol via Codex.
R20 follow-up Give inspector reasoning markdown its projected source thread and retain the explicit environment fallback for proposed plans without a thread reference. Workspace links and images now resolve through their owning environment after removal of the active-environment fallback. Implemented by GPT-5.6 Sol via Codex.
Read the inclusive cursor, a full history page, and a look-behind row so older history does not terminate after one page. Finding: P01 pagination termination Model: GPT-5.6 Sol via Codex
Cancel pending SDK requests before aborting the native session. Preserve per-admission cancellation state and treat stopped initial prompts as interruption instead of provider failure. Finding: R14 prompt cancellation Model: GPT-5.6 Sol via Codex
Keep the original cursor owner through ancestor traversal and preserve the history budget across empty intermediate forks. Verify exact paged history against the complete nested projection. Finding: P01 nested lineage Model: GPT-5.6 Sol via Codex
Retain pending admission after transient status failures and use one generation-owned retry worker. Ignore stale timers and duplicate evidence so older prompts cannot finish newer steering. Finding: R14 status reconciliation Model: GPT-5.6 Sol via Codex
Keep hidden local and inherited rows from consuming history pages. Preserve stop-request dependencies, source-run cutoffs, and imported history while loading related metadata from the selected cohort and using indexed watermark lookups. Finding: P01 bounded history visibility Model: GPT-5.6 Sol via Codex
Adapt grouped tool summaries and the floating working timer to V2 run, attempt, and queue state. Bring over the composer, keyboard, and disclosure transitions while retaining the V2 activity inspector and queue controls. Keep OV2 web composer and grouping behavior intact; share only the existing command label parser with mobile.
Restores main features dropped by the policy replay: #8569 theme wiring, settings search rework, #8803 workspace-mutation refresh (v2-adapted), video + image previews (web and mobile, v2-adapted), #8862 Expo glass, and the round's docs. Timeline thinking rows (#8984) stay on the v2 work-live system. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The v2 equivalents of main's #8984 and #8922: a "Working for ..." header anchors the active run, the trailing live tool row survives between actions in past tense instead of vanishing, and a shimmering Thinking row marks reasoning gaps. During workspace preparation the header shows "Setting up worktree..." (driven by the local dispatch flag or the v2 run's preparing status, so remote viewers see it too), the composer footer span is gone, and draft promotion waits until the run starts or startup fails instead of navigating mid-preparation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adopts the round's main features into the v2 architecture: the #9023 media rework (streamed videos, media-file assets, protocol-relative links), #9098 shared live-activity row folded into the v2 working and thinking rows, the #9084/#9078 Claude model catalog for v2 consumers, a native #9005 OpenCode child-session abort in the v2 adapter, #9013's landed LegendList patch, and per-environment sidebar provider entries. For #8600 the server-side pieces land, but auto-settle evaluation stays client-side (reading the new server-owned settings) until the v2 orchestrator grows its own settlement reactor; main's v1-only reactor and coalescer additions are dropped with the rest of the v1 path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ator Ports #8600's server-owned settlement to orchestration v2 instead of keeping client-side evaluation. A ThreadSettlementService sweep runs at startup, on auto-settle settings changes, and once per minute: it evaluates inactivity and merged or closed pull requests over v2 thread shells and dispatches the new guarded thread.auto-settle command, which rejects threads that changed after the sweep's snapshot or carry any explicit override, then reuses the orchestrator's settle lifecycle. With the server deciding, the clients drop their effectiveSettled evaluation and partition on the persisted settledOverride like main. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adopts main's round-19 features into the v2 stack: payload-budgeted orchestration replay (#8992), sidebar row subscription leases (#9052), tool group virtualization and scroll anchoring (#9106), repeated-command and browser-group presentation, inline assistant citations (#9146), per-cwd provider skills discovery (#8778), Claude composer skill dispatch (#9128), grok health probe and model negotiation (#9154), and the failed-tool thinking fallback (#9165). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Round 17 adopted main's #8850 ComposerBanner.Attachment (mx-auto plus the standalone drawer-inset width) without main's matching mounts, so the stash tab's ml-auto lost to the attachment's auto right margin and the tab centered over the composer. Column now spans its attachments like main does, the stash tab zeroes the right margin, and the stash menu keeps the full dock width. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Orchestrator V2 already projects native subagents, and the existing Agents panel already owns their live detail. The remaining problem is presentation: chat repeats one lifecycle row per agent, duplicating the roster and stretching the turn.
This PR keeps the existing orchestration semantics and:
There are no server, contract, persistence, provider-adapter, recovery, client-runtime, or mobile behavior changes.
Screenshots
Verification
Built and tested with GPT-5.6-Sol in the Codex harness through T3 Code.
Note
Low Risk
Web-only chat timeline presentation; behavior changes only when
onOpenAgentsis wired, with no auth, API, or persistence impact.Overview
Local subagent events in chat are collapsed into one summary row per run (or per attempt when steered), instead of repeating a lifecycle card per agent.
deriveMessagesTimelineRowsgains optionalsummarizeSubagentsand a newagent-spawnrow type with aggregated counts (working, failed, stopped). Inherited subagents and runs without a summarization callback still render as ordinary event rows; superseded vs active attempts stay in separate groups.MessagesTimelinetakesonOpenAgentsand enables summarization only when that callback is present. The newAgentSpawnCtaRowshows status text (including visible failures while others are still running) and opens the existing Agents panel.ChatViewpassesaddAgentsSurfacefor that hook.Tests cover grouping, attempt boundaries, opt-out when summarization is off, and the updated markup expectations.
Reviewed by Cursor Bugbot for commit 28a337d. Bugbot is set up for automated code reviews on this repo. Configure here.
Note
Add orchestration agent summary rows to
MessagesTimelineagent-spawnrow variant inderiveMessagesTimelineRowsthat collapses local subagent events into a single aggregated summary per attempt, run, or item whensummarizeSubagentsis enabled.AgentSpawnCtaRowcomponent, which renders a clickable CTA with dynamic status text (live vs completed, failed/stopped counts) and opens the Agents panel viaonOpenAgents.MessagesTimelineaccepts an optionalonOpenAgentsprop; when present it enables subagent summarization and threads the callback throughTimelineRowCtx.ChatViewContentwiresonOpenAgents={addAgentsSurface}to the timeline so users can open the Agents panel from the summary row.isRowUnchangedextended to shallow-compareagent-spawnrows, preventing unnecessary re-renders in virtualization.onOpenAgentsis provided, local subagent event rows are no longer rendered individually — callers relying on per-event row rendering (e.g. in MessagesTimeline.logic.ts) should verify the collapsed summary covers their use case.Macroscope summarized 28a337d.