fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(client-runtime): recover subscriptions that lose their transport - #4602

Open
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery
Open

fix(client-runtime): recover subscriptions that lose their transport#4602
colonelpanic8 wants to merge 2 commits into
pingdotgg:mainfrom
colonelpanic8:t3code/subscription-transport-recovery

Conversation

@colonelpanic8

@colonelpanic8colonelpanic8 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

A durable subscription that fails with a transport-level RpcClientError now asks the supervisor to re-establish the session, instead of draining and waiting for a session that may never come.

Requests are coalesced per session via a WeakSet keyed on the session object, so a burst of subscription failures produces exactly one reconnect request.

Why

subscribeDynamic handles a transport failure by logging and draining, deferring to "the next session". Nothing guarantees one arrives. supervisor.session is written only when a connection lease is established (connection/supervisor.ts:532) or cleared on teardown (:243), and the only liveness probe runs on an application-active wakeup — a desktop window that stays visible never health-checks its connection.

So a stream that dies while the session stays usable — a half-open socket (NAT idle timeout, sleep/resume, VPN flap, where no close is ever delivered), or a failure confined to that one stream — leaves the subscription dead for the lifetime of the session.

Nothing surfaces it. Unary RPCs on that same session keep working, so the app looks connected and commands still round-trip while the environment's state silently stops updating.

This was diagnosed from a case where client and server were on different machines, so both sides could be inspected independently (#4589). On one connection at one moment:

  • ten thread.settle commands were accepted server-side, with receipts
  • the pre-existing shell subscription was dead — a new thread created during the incident never appeared in the sidebar, though the server had it running
  • a thread-detail subscription created after the failure worked fine

That asymmetry is the signature: the drain is per subscription stream, so subscribeToSession() runs fresh for anything started afterward while already-drained streams stay dead. A connection-level failure could not produce it — it would take out the new subscription too. Other clients against the same server were unaffected throughout.

The user-visible result is an environment frozen in place: threads stuck showing Working long after they finished, new threads missing from the sidebar, and lifecycle actions succeeding as no-ops against a stale view. Only restarting the client recovers it, because a restart is the one thing that makes a new session.

Note that the branch immediately below this one already retries within the same session when retryExpectedFailureAfter is set (state/shell.ts passes "250 millis"). Only the transport branch had no path back.

Why coalesce

An environment carries many durable subscriptions — the shell, the thread list, every open thread — and a dead transport tends to take all of them at the same moment. Each retryNow is a signal the supervisor acts on, and a signal delivered mid-establishment aborts that attempt (supervisor.ts:362-384). N unfiltered requests would abort N reconnects and stall recovery exactly when it is needed. The regression test for this fails with expected 2 to be 1 if the guard is removed.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes — no UI change
  • I included a video for animation/interaction changes — n/a

One existing test changed its expectation. "keeps durable subscriptions alive across a transport failure and new session" asserted retryCount === 0, and passed only because the test manually supplies the replacement session that production has no way to produce. It now expects one request.

Validation (scoped per AGENTS.md, which reserves repo-wide vp check / vp run typecheck for CI):

vitest run packages/client-runtime/src/ # 473 passed
tsgo --noEmit # packages/client-runtime — clean
vp lint --report-unused-disable-directives # no findings in changed files

Not addressed here, both worth their own change: unary request (rpc/client.ts:107) has no timeout, so a call into a hung session never settles — which can latch an in-flight guard in the UI; and there is no periodic liveness probe, only the application-active one.


Note

Medium Risk
Changes connection recovery for all durable environment subscriptions; incorrect coalescing or stale-session handling could over- or under-reconnect, but scope is limited to transport failures and is heavily tested.

Overview
Durable RPC subscriptions that fail with a transport-level RpcClientError no longer only log and drain while hoping the supervisor swaps sessions. subscribeDynamic now calls requestSessionRecovery, which triggers supervisor.retryNow when the failing session is still the active one, so reconnect can happen even when the socket looks fine to unary calls.

Recovery is coalesced per session with WeakSet guards so many subscriptions dying together produce one reconnect signal; stale sessions and in-flight duplicates are ignored. Domain failures still do not reconnect.

Tests cover single and multi-subscription reconnect, stale-session filtering, resubscribe interrupting recovery, and the updated expectation that transport failure implies one retryNow.

Reviewed by Cursor Bugbot for commit b5836f1. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix subscribeDynamic to recover subscriptions that lose their transport

  • When an RpcClient subscription fails with a transport error, subscribeDynamic now calls supervisor.retryNow before draining to await the next session.
  • A new requestSessionRecovery effect coalesces recovery signals per session: only one retryNow is issued per broken session, and signals from already-replaced sessions are ignored.
  • New tests in client.test.ts cover deduplication, stale-session filtering, and in-flight resubscription interleaving.
  • Behavioral Change: subscriptions that previously silently drained on transport loss now actively trigger supervisor reconnection.

Macroscope summarized b5836f1.

@coderabbitai

coderabbitaiBot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 38f0bf65-d02c-404b-adf3-0c79451f3d23

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Jul 27, 2026
Comment threadpackages/client-runtime/src/rpc/client.ts Outdated

@cursorcursorBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit b80214c. Configure here.

Comment threadpackages/client-runtime/src/rpc/client.ts Outdated
@macroscopeapp

macroscopeappBot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces new subscription recovery behavior in the client-runtime connection layer, including module-level state tracking and changes to how subscriptions coordinate with the supervisor during transport failures. While well-tested, these are meaningful changes to connection lifecycle management that warrant human review.

You can customize Macroscope's approvability policy. Learn more.

colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
Adds the root-cause fix for the frozen-environment bug: a durable
subscription that dies with a transport error now asks the supervisor to
re-establish, instead of draining and waiting for a session nothing
produces. Placed next to pingdotgg#4405, the only other topic touching
rpc/client.ts and subscription lifecycle.
Also repins pingdotgg#4593 to 6980a8c, which stops an immediate resync from
leaving an obsolete trailing timer scheduled.
Rebuilt with --mode reproduce; 12 conflicts, all replayed verbatim from
the prior build with nothing missing. Syntax gate clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from b80214c to c3cb473CompareJuly 27, 2026 11:14
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 27, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from c3cb473 to 5b840c4CompareJuly 28, 2026 02:08
colonelpanic8and others added 2 commits July 27, 2026 19:47
A durable subscription that fails with a transport-level RpcClientError
drains and defers to "the next session". Nothing guarantees one arrives:
the supervisor replaces a session only when the socket reports closed or
an `application-active` probe fails, and there is no periodic liveness
check. A stream that dies while the session stays usable — a half-open
socket, or a failure confined to that stream — therefore leaves the
subscription dead for the lifetime of the session.
Nothing surfaces it. Unary calls on that same session keep working, so
commands still round-trip while the environment's state silently stops
updating: threads frozen mid-turn, new threads absent from the sidebar,
lifecycle actions succeeding as no-ops against a stale view. Only
restarting the client recovers it, because a restart is the one thing
that makes a new session.
Ask the supervisor to re-establish instead, so the drain hands off rather
than dead-ending. Requests are coalesced per session: a dead transport
takes every subscription for an environment at once, and each `retryNow`
aborts an in-flight establishment attempt, so asking once per
subscription would stall the reconnect it is trying to cause.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@colonelpanic8
colonelpanic8force-pushed the t3code/subscription-transport-recovery branch from 5b840c4 to b5836f1CompareJuly 28, 2026 02:51
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Jul 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 10, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 11, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 18, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 28, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Aug 31, 2026
colonelpanic8 added a commit to colonelpanic8/t3code that referenced this pull request Sep 2, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@colonelpanic8