fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(server): back off relay client restarts - #8427

Closed
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff
Closed

fix(server): back off relay client restarts#8427
Adolanium wants to merge 1 commit into
pingdotgg:mainfrom
Adolanium:fix/relay-client-restart-backoff

Conversation

@Adolanium

@AdolaniumAdolanium commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

What Changed

The Cloudflare relay supervisor now waits before it respawns cloudflared.

If the child exits and that config is still desired, the next start waits 500ms, then 1s, 2s, up to 10s. The wait does not hold the reconcile lock, so a new applyConfig still starts right away and resets the backoff.

Why

A bad tunnel token or a child that dies on start was a tight spawn loop. CPU plus process churn, with TUNNEL_TOKEN in the child env the whole time.

Native telemetry already backs off the same way. The relay client did not.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included before/after screenshots for any UI changes (N/A)
  • I included a video for animation/interaction changes (N/A)

Note

Medium Risk
Changes connector supervision and timing for cloud tunnel availability; mis-tuned backoff or lock handling could delay recovery or allow brief spawn races, but scope is limited to managed relay restarts.

Overview
Adds exponential backoff when the Cloudflare relay (cloudflared) supervisor respawns a child that exited while the same tunnel config is still desired. Delays start at 500ms, double per failure, and cap at 10s via exported relayClientRestartDelay.

The supervisor releases the reconcile lock during the sleep, then re-acquires it to call reconcile only if nothing else is active and the desired config still matches. applyConfig resets the attempt counter to 0, so an explicit config change can restart immediately instead of waiting through backoff.

Tests cover the delay curve and use TestClock so the post-exit restart test asserts no second spawn until ~500ms has passed.

Reviewed by Cursor Bugbot for commit 90f9ea0. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add exponential backoff to relay client restarts in ManagedEndpointRuntime

Relay client restarts now delay with exponential backoff starting at 500 ms, doubling per attempt, and capped at 10 seconds via relayClientRestartDelay. A new restartAttemptRef tracks the attempt count; applyConfig resets it to 0 whenever a new config is applied. After the backoff sleep, superviseConnector re-checks state before restarting so it only restarts if the connector is still desired and not already active. Behavioral Change: restarts are no longer immediate — there is now a delay of up to 10 seconds between process exit and respawn; existing callers expecting instant restarts will see delayed restarts.

Macroscope summarized 90f9ea0.

When cloudflared exited, the supervisor spawned it again immediately. A bad tunnel token or a child that dies on start was a tight spawn loop.
Wait 500ms, then 1s, 2s, up to 10s, before the next start. A new applyConfig still starts right away.
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 11a3d583-7068-4e64-b39e-0603a6379e02

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@github-actionsgithub-actionsBot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 27, 2026
@macroscopeapp

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 90f9ea0

Macroscope's review found this PR approvable — This focused server-side fix limits repeated Cloudflare relay respawns with a tested exponential backoff, while preserving immediate explicit configuration reconciliation and guarding against stale restarts. Its production impact is confined to connector recovery after process exits, with no schema, deployment, or sensitive-package changes.

You can add or adjust custom eligibility rules. Learn more.

@t3dotgg

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

We're closing this PR as we clean up the T3 Code backlog. Thank you for taking the time to put this together.

We are keeping the relay restart fix in the open PR #8788 instead of maintaining two implementations. It addresses the same rapid-exit loop with capped exponential delays, resets after stable uptime or a configuration change, and keeps configuration changes responsive while it waits. The fix is not on main yet.

If you believe we closed this in error, please reopen the PR and leave a comment explaining what we missed.

@t3dotggt3dotgg closed this Sep 1, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M30-99 changed lines (additions + deletions).vouch:unvouchedPR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Adolanium@t3dotgg