fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(client): a busy backend no longer looks like a disconnect - #7233

Open
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout
Open

fix(client): a busy backend no longer looks like a disconnect#7233
btsouth wants to merge 1 commit into
pingdotgg:mainfrom
btsouth:fix/tolerate-transient-connection-probe-timeout

Conversation

@btsouth

@btsouthbtsouth commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Fixes#7231.

Problem

When the app comes to the foreground, the supervisor probes the live session with a 15s deadline. A probe that misses that deadline was treated exactly like a definite failure: the lease was replaced and the environment dropped to "Failed to connect. Reconnecting...".

The socket was still open and no close event had arrived. The backend was busy, not gone. The visible cost is the composer, which is disabled while the client recovers, and it lands precisely when someone has returned to the app to send a message.

Fix

On desktop and web, a missed deadline now waits 5s and probes once more. Only a second miss replaces the lease.

The retry lives inside the forked probe effect rather than in the signal loop, which matters for two reasons:

  • Disconnect, explicit retry, going offline, and resume signals still interrupt it exactly as before, through the existing Fiber.interrupt(probe) paths.
  • Recovery stays bounded. An earlier version of this fix tolerated a timeout and waited for the nextapplication-active wakeup to decide. On desktop that wakeup only fires on visibilitychange, so a user who stays in the window would never produce one, and a backend wedged at the RPC layer (socket healthy, pings answered, handler deadlocked) would have sat at "connected" indefinitely with every request hanging. Doing the retry inline caps that at roughly 35s instead.

Switching the desktop path to Effect.timeoutOption also removes the guesswork about whether a given failure was our own deadline: a missed deadline arrives as None on the success channel, and every real probe failure still fails immediately. That matches the idiom already used by the provider probes on the server side.

Mobile is unchanged.application-active-probe keeps its 3s fail-fast and does not retry. After a background suspension the OS has usually killed the socket for real, so waiting longer would only delay recovery. Mobile cannot emit the bare application-active reason, so the retry path is unreachable from it.

Worth noting for reviewers: a genuinely dead transport is still detected independently of this path by the RPC pinger through session.closed, so nothing here is the sole defence against a dropped socket.

Relationship to #3553, #4137, and #5198

#3553 reported this and was closed by #4137, which made the probe itself lightweight (serverProbe rather than serverGetConfig). That was the right fix for probe cost. It did not change the teardown rule, and on a host that is starved rather than merely doing extra work a cheap probe misses its deadline too.

#5198 proposed tolerating a transient timeout and had the right instinct. It has carried merge conflicts since 1 Aug. This PR takes the same starting point and adds the retry, without which tolerance alone can strand a wedged backend.

Backend starvation itself is tracked in #4773. This is the client's reaction to it, which is worth fixing separately because the client cannot assume the server will ever be fast.

Testing

  • Reworked the existing "stalled desktop foreground probe" test, which asserted teardown on the first timeout. Its 14999ms / 1ms deadline-boundary assertions are preserved on the retry.
  • Added a test that the retry answering keeps the session, the generation, and the composer path untouched.
  • 36/36 supervisor tests pass, over 5 consecutive runs.
  • typecheck and vp lint packages/client-runtime/src/connection clean.

Docs: docs/internals/connection-runtime.md "Wakeups" section now states the retry, the bound, and the mobile exclusion.

Model: Claude Opus 5; harness: Claude Code.


Note

Medium Risk
Changes connection supervisor wake/probe policy on desktop/web, which affects when sessions are replaced and the composer is disabled; behavior is well-tested but touches core connectivity logic.

Overview
Desktop/web foreground health checks no longer tear down the live session on a single 15s probe timeout. A miss is treated as a busy backend (socket still open); the supervisor waits 5s, runs one more probe, and only then fails and replaces the lease. Real probe failures still reconnect immediately; mobile application-active-probe stays 3s fail-fast with no retry.

monitorConnectedLease switches the desktop path to Effect.timeoutOption and runs the retry inside the forked probe effect so disconnect/offline/resume signals still interrupt via existing Fiber.interrupt paths, with worst-case recovery bounded (~35s).

Docs and supervisor tests cover retry success (stay connected), double stall (reconnect), and preserved mobile behavior.

Reviewed by Cursor Bugbot for commit 1164990. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Fix desktop/web foreground probe to retry once before treating a stall as a disconnect

  • A single missed 15s probe on desktop/web no longer immediately reconnects; the supervisor now waits 5s (CONNECTION_PROBE_RETRY_DELAY) and issues a second probe before treating the timeout as a failure.
  • Mobile (application-active-probe) retains the original behavior: 3s deadline, no retry, immediate failure on timeout.
  • Behavioral Change: desktop/web clients on a slow/busy backend will stay in the current session instead of re-establishing a new lease on the first stalled probe.

Macroscope summarized 1164990.

A foreground health probe that missed its 15s deadline was treated the
same as a definite failure, so one slow answer replaced a live session
and put the environment into "Reconnecting". The socket was still open
and no close had arrived; the backend was busy, not gone. While the
client recovered, the composer was disabled, which is what users
actually feel, and it lands exactly when someone has come back to the
app to send a message.
Desktop and web foreground probes now wait 5s after a missed deadline
and ask once more. Only a second miss replaces the lease, which bounds
recovery for a genuinely wedged backend at roughly 35s instead of
leaving it to the next foreground wakeup that may never come. The retry
lives inside the forked probe effect, so disconnect, retry, offline, and
resume signals still interrupt it exactly as before.
Using timeoutOption for the desktop path also removes the guesswork
about which failure was our own deadline: a missed deadline arrives as
None, and every real probe failure still fails immediately.
Mobile resume probes are unchanged. After a background suspension the
socket usually is dead, so the 3s fail-fast is correct there.
Reworks an existing test that asserted the old teardown-on-first-timeout
behavior, keeping its deadline-boundary assertions, and adds coverage
for the retry-answers path.
Model: Claude Opus 5; harness: Claude Code.
@github-actionsgithub-actionsBot added the vouch:unvouched PR author is not yet trusted in the VOUCHED list. label Aug 16, 2026
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 66102392-47da-4c11-82a2-8373fe7003f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added the size:M 30-99 changed lines (additions + deletions). label Aug 16, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR changes runtime connection probe behavior - desktop/web probes now retry once before failing on timeout. Connection supervision is core reliability infrastructure and the behavioral change warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: One slow foreground health probe tears down a live session and disables the composer

1 participant

@btsouth