fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863) - #871

Merged
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race
Aug 27, 2026
Merged

fix(ci): wait for the metrics.k8s.io APIService before the e2e helm install (#863)#871
aptracebloc merged 1 commit into
developfrom
fix/863-seal-check-metrics-apiservice-race

Conversation

@aptracebloc

@aptraceblocaptracebloc commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes#863.

The flake

Seal-check egress-enforcement (k3d) failed intermittently before it tested anything:

resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered.

This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight lookups (client#823) — asynchronously, after nodes report Ready. The e2e harnesses gated only on kubectl wait --for=condition=Ready nodes and then helm-installed, so on a fast runner the install beat the addon and the preflight failed the whole release. Evidence: #862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts.

The fix

e2e_wait_for_metrics_apiservice (new, in scripts/tests/lib/e2e-common.sh), called after node-Ready and before the helm install in every harness that installs a preflight-carrying chart directly:

  • e2e-seal-check.sh — the reported flake (local chart).
  • e2e-full-seal.sh — its secret-ful sibling (local chart).
  • e2e-auto-upgrade.shsame class: its first install is the last published chart, which carries the preflight too (fix-the-class, caught in self-review).

The helper polls for the APIService to EXIST first, then best-effort waits for Available. This matters: the object itself appears late (not merely its condition), and kubectl wait on a not-yet-created named object errors NotFound rather than waiting — so a bare kubectl wait --for=condition=Available apiservice/... (the issue's first-draft one-liner) would just swap one red for another in the same window. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race.

Rejected the weaker options (resourceMonitor: false / nodeAgents.metricsServerPreflight: false): both go green only by deleting the client#823 preflight coverage this seal-check exists to exercise on a real cluster.

e2e-metrics-apiservice-wait.bats (new) pins the invariant — all three harnesses wait before their first install, and the wait polls-for-existence before the condition wait — so the guard cannot silently drift back out.

Verification

Verified on real k3d (rancher/k3s:v1.36.3-k3s1):

  • Reproduced the exact bug: with the APIService absent, helm install --dry-run=server fails at client/templates/resource-monitor-daemonset.yaml:69 with the issue's error.
  • Confirmed the fix: the helper blocks until the APIService is registered + Available, after which the preflight renders tracebloc.io/metrics-server-preflight: satisfied-by-apiservice.
  • Confirmed the helper's loud-fail path (clear "cluster bring-up problem, not a chart defect" message) and correct behavior under set -euo pipefail.
  • make lint clean; full bats suite failure set identical to develop tip (zero failures added); new guard test green.

Note

Low Risk
Changes only CI e2e bring-up scripts and static bats guards; no production install or chart template behavior is modified.

Overview
Fixes intermittent Seal-check egress-enforcement (k3d) failures where Helm aborted before any seal test with metrics.k8s.io/v1beta1 API is not registered — a harness race because k3s registers the metrics APIService after nodes go Ready, while the chart’s resource-monitor preflight requires it at install time.

Adds e2e_wait_for_metrics_apiservice in e2e-common.sh: poll until v1beta1.metrics.k8s.io exists (avoiding a bare kubectl waitNotFound on a late-created object), then a best-effort Available wait. The three harnesses that install preflight-carrying charts call it after node-Ready and before their first helm install: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (published chart path included).

e2e-metrics-apiservice-wait.bats locks in that contract so the wait cannot drift out of any harness or revert to existence-less kubectl wait.

Reviewed by Cursor Bugbot for commit 89f19f6. Bugbot is set up for automated code reviews on this repo. Configure here.

…nstall (#863)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@aptraceblocaptracebloc self-assigned this Aug 27, 2026

@saadqbalsaadqbal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diagnosis is the valuable half. #862 false-failing at 26s while #861 passed at 51s, neither touching the chart or these scripts, is a timing signature rather than a hunch — and kubectl wait on a not-yet-created named object errors NotFound rather than waiting is the non-obvious fact that separates fixing this from re-reddening it with a different message. The issue's first-draft one-liner would have done exactly that, and you said so.

Polling for existence and then best-effort waiting for Available is right, and mirroring the production installer's _wait_for_metrics_apiservice rather than inventing a second shape is right too. So is refusing resourceMonitor: false / metricsServerPreflight: false — both go green by deleting the client#823 coverage this seal-check exists to exercise, which is going green by removing the thing under test.

Both bats assertions are genuinely derived, which I checked: call_ln < install_ln from real line numbers, get_ln < wait_ln from the awk-extracted function body, and [ -n … ] guards that fail when a grep finds nothing — including "no 'helm install' in $f — test assumption broken", the test checking its own premise. The call-line regex being start-of-line anchored also means a mention in a comment can't satisfy it.

One finding, non-blocking: the assertions are derived but the scope is a hardcoded list.HARNESSES names three scripts. A fourth harness that helm-installs a preflight-carrying chart is simply invisible to the guard — it passes by not being enumerated.

The evidence that this is the part that misses things is in your own PR: e2e-auto-upgrade.sh was "same class, caught in self-review". Your first pass at the list missed one of three, and the guard cannot notice the next one.

Two sibling guards in this repo took the other road within the same week — client#860's CronJob check reads every CronJob out of the rendered manifests and holds no list of names (refusing outright on zero), and backend#2655 derives its terraform test directories from the tree. Deriving here looks tractable: the exclusion reason you give — the e2e-cluster/journey/proxy harnesses stop before the tracebloc install, so they render no resource-monitor and carry no preflight — is itself a property of the script (does it helm install the tracebloc chart at all?) rather than a fact only a human knows. Worth doing when the fourth harness appears, if not now.

Green, no threads, verified on real k3d. 👍

@aptracebloc
aptracebloc merged commit 1306009 into developAug 27, 2026
52 of 53 checks passed
@aptracebloc
aptracebloc deleted the fix/863-seal-check-metrics-apiservice-race branch August 27, 2026 11:31
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ci): seal-check helm-installs before k3s registers the metrics-server APIService

3 participants

@aptracebloc@saadqbal@LukasWodka