Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Wait for the load balancer address instead of failing on the first read - #53

Merged
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address
Aug 19, 2026
Merged

Wait for the load balancer address instead of failing on the first read#53
kierenj merged 1 commit into
developfrom
feat/wait-for-ingress-address

Conversation

@kierenj

Copy link
Copy Markdown
Member

The problem

create reads the ingress address once, immediately after the rollout it waits for, and gives up if it is not there yet:

getting external IP...
jq: error (at <stdin>:54): Cannot iterate over null (null)
failed

The address is published by the ingress controller on its own cycle, independently of — and later than — the deployment rollout. So that read is usually too early. The job fails, Kubernetes retries it, and a later attempt eventually catches it. The Red Pepper MCP deploy hit this three times in dev and once in live; it only ever looked like flakiness.

How long it actually takes

Measured on both clusters by creating an ingress with a fresh hostname and timing until .status.loadBalancer.ingress[0].ip appeared:

clusternminmedianmax
NonLive619.1s51.8s56.0s
Live438.3s51.5s55.4s

The scattered minima against a hard ceiling near 56s is the signature of a single publication cycle — the job arrives at a random point in it, not variable load. Both clusters behave identically.

The fix

A bounded wait, polling until the address appears.

ADDRESS_TIMEOUT defaults to 120s and ADDRESS_POLL_SECONDS to 5s.

120s is deliberately more than 1.5× the measurement (~84s). Because the mechanism is a discrete ~55s cycle rather than a continuous distribution, ~84s covers barely more than one cycle and would fail outright if a single publication were missed — a controller restart or leader-election handover. 120s survives that. The asymmetry favours it: too generous only delays reporting a genuine failure, too tight reverts to spurious deploy failures. Both are overridable.

On failing fast instead

There is nothing to check. IngressStatus is only:

status.loadBalancer.ingress[] → { hostname, ip, ports[{ error, port, protocol }] }

No conditions, no phase. A pending address and a permanently broken one are byte-identical — both an empty loadBalancer. (ports[].error exists but ingress-nginx never populates it.) So the timeout prints the object's events, where a real failure does show up, and exits 1.

Verification

caseresult
address appears on the 3rd pollwaits, then proceeds
address never appearstimes out, prints events, exit 1

Record-matching regression matrix from #52, all unchanged:

create pepper exit=0 PATCH .../AAA
create pepper-mcp exit=0 PATCH .../BBB
delete pepper exit=0 DELETE .../AAA
create brand-new exit=0 create path
delete missing exit=0 no request
create dupe exit=1 refusing to guess
delete dupe exit=1 refusing to guess

Also collapses the duplicated ingress/service and ip/hostname branches, which held four copies of the same fetch-and-check.

Rollout

Needs 0.0.25 published, then a chart bump. k8s#126 currently pins 0.0.24 — if that has not merged yet it is cheaper to retarget it at 0.0.25 and ship both in one bump.

🤖 Generated with Claude Code

The create path read the ingress address once, immediately after the rollout it waits for, and gave
up if it was not there yet - "jq: error ... Cannot iterate over null" followed by "failed". The
address is published by the ingress controller on a separate cycle, so that read is usually too
early: the job failed and was retried until a later attempt happened to catch it. The Red Pepper MCP
deploy failed three times this way in dev and once in live.
Measured on both clusters by creating an ingress with a fresh hostname and timing the address:
NonLive n=6 min 19.1s median 51.8s max 56.0s
Live n=4 min 38.3s median 51.5s max 55.4s
The hard ceiling near 56s with scattered minima is one publication cycle - the job arrives at a
random point in it. ADDRESS_TIMEOUT therefore defaults to 120s, two cycles, so a single missed one
is survivable; ADDRESS_POLL_SECONDS defaults to 5.
On timeout the object's events are printed and the job exits 1. There is no state to check instead:
IngressStatus carries only loadBalancer.ingress[] and no conditions, so a pending address and a
permanently broken one are indistinguishable except by waiting.
Also collapses the duplicated ingress/service and ip/hostname branches, which had four copies of the
same fetch-and-check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CopilotAI lite review requested due to automatic review settings August 19, 2026 06:49

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR addresses flaky create runs by waiting for Kubernetes to publish the external load balancer address (IP/hostname) instead of reading it once immediately after rollout and failing when the address is not yet present.

Changes:

  • Add a bounded polling loop (with ADDRESS_TIMEOUT and ADDRESS_POLL_SECONDS) to wait for the ingress/service external address before writing the DNS record.
  • Reduce duplicated branches by unifying ingress vs service selection and IP vs hostname selection.
  • Document the new waiting behavior and configuration knobs in the README, and bump the script version to v0.0.25.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

FileDescription
readme.mdDocuments the new bounded wait for ingress/service external address and the related environment variables.
k8s-tools.shImplements polling for the external address with configurable timeout/interval; bumps version string.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadk8s-tools.sh
Comment on lines +99 to +103
resource=$(kubectl --namespace=$namespace get $kind $target --output=json 2>/dev/null)
if [ -n "$resource" ]; then
dns_record_value=$(echo "$resource" | jq -r ".status.loadBalancer.ingress[0].$field // empty")
[ -n "$dns_record_value" ] && break
fi
@kierenj
kierenj merged commit 684bbcf into developAug 19, 2026
1 check passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@kierenj