fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(cloud): back off relay connector restarts - #8906

Closed
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff
Closed

fix(cloud): back off relay connector restarts#8906
nateEc wants to merge 1 commit into
pingdotgg:mainfrom
nateEc:codex/fix-8760-relay-restart-backoff

Conversation

@nateEc

@nateEcnateEc commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Fixes#8760.

An immediately failing relay child was restarted in a tight loop, creating processes and trace records until the server exhausted its heap.

Wait five seconds before retrying a connector exit. The wait intentionally releases the configuration lock so users can disable or change T3 Connect immediately.

Verification: vp fmt --check, focused managed endpoint runtime tests (8 passing), and server typecheck (existing suggestions only).

Model and harness: GPT-5 Codex via Codex CLI.


Note

Medium Risk
Changes connector supervision and config locking timing for cloud relay processes; mis-timed restarts or lock handling could affect T3 Connect availability or responsiveness to config changes.

Overview
Fixes a tight restart loop when the Cloudflare relay child exits immediately, which could spawn processes and trace records until the server ran out of heap.

superviseConnector now waits 5 seconds (RELAY_CONNECTOR_RESTART_DELAY) before calling reconcile again. The delay runs outside the reconcile semaphore so applyConfig can still disable or change T3 Connect without blocking behind a failing child. After the wait, it re-reads desiredConfig and only restarts if the same tunnel config is still wanted.

The supervisor test was renamed and updated to use TestClock: it asserts no second spawn for the first 4 seconds after exit, then restart after the full 5 second delay.

Reviewed by Cursor Bugbot for commit 14468ba. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Delay relay connector restarts by 5 seconds in ManagedEndpointRuntime

Introduces RELAY_CONNECTOR_RESTART_DELAY (5 seconds) in ManagedEndpointRuntime.ts. After a managed Cloudflare connector exits, the runtime now sleeps before attempting a restart and releases the reconcile semaphore during the wait, so other operations are not blocked. After the delay, it re-validates that the desired config still matches the exited connector before calling reconcgeConfig; if the config changed, the restart is skipped. Updates the supervision test to use TestClock and assert the 5-second wait with no immediate respawn.

  • Behavioral Change: restarts of exited relay connectors are now delayed by 5 seconds and gated on a second config validation; callers expecting immediate respawn will observe the delay. The reconcile semaphore is released during the delay, opening a window where config changes can cancel the restart.

Macroscope summarized 14468ba.

Wait before restarting an exited relay connector to prevent rapid crash loops from exhausting the host.
Release the configuration lock during the delay so T3 Connect can still be changed or disabled immediately.
Verify with focused managed endpoint runtime tests and server typecheck.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 44501e4f-4a56-49f1-9e3c-2cb4af9c8bf3

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:S 10-29 changed lines (additions + deletions). labels Aug 31, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — Managed Cloudflare connector restarts now use a fixed five-second delay on an existing production path, with configuration changes able to cancel or supersede the retry. The change is focused and test-covered, but the new default recovery policy warrants human review.

You can add or adjust custom eligibility rules. Learn more.

@derektrimm

Copy link
Copy Markdown
Contributor

Heads-up on overlap: #8788 (opened Aug 30) addresses the same issue (#8760) in the same supervisor. It uses exponential backoff (1s doubling to a 60s cap) with an immediate first restart, resets after 30 seconds of stable uptime or an explicit config change, and keeps the delay outside the reconcile semaphore so a user disabling or changing T3 Connect is never blocked behind it. It ships TestClock tests for the doubling, the stable-uptime reset, and config-change preemption, and Macroscope approved it at head. Flagging so the maintainers can compare the two directly rather than review them independently.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It covers the same failing-connector loop as this five-second delay, with capped exponential delays and a reset after stable uptime or a configuration change. Both let configuration changes proceed during the wait. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S10-29 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Failing cloudflared child respawns without backoff until the server exhausts its V8 heap

3 participants

@nateEc@derektrimm@t3dotgg