docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

docs(install): write down the release-name convention, and guard it (backend#2621) - #861

Merged
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention
Aug 27, 2026
Merged

docs(install): write down the release-name convention, and guard it (backend#2621)#861
LukasWodka merged 3 commits into
developfrom
feat/2621-release-name-convention

Conversation

@LukasWodka

@LukasWodkaLukasWodka commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Closes tracebloc/backend#2621

What was actually wrong

The installer has always named the release after the namespace:

# scripts/lib/install-client-helm.sh:2272
helm upgrade --install "$TB_NAMESPACE""$chart_ref" --namespace "$TB_NAMESPACE" --create-namespace

…defaulting both to tracebloc, and it explains itself at the point of decision (line 1862):

"The on-prem client is one-per-machine and is identified to the backend by its credentials (clientId), not by this name — so we don't ask the user to invent one."

So the self-service path is consistent by construction. The defect was that this was written down nowhere. docs/INSTALL.md's only example was helm install my-tracebloc …, and docs/migration-tools/tenant-config.example.env records the consequence as a fact to work around rather than a defect to fix:

"The Helm release name (NOT always the namespace name; tenant-a's release is tracebloc)"

Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.

The case that prompted it: an engineer's given name ends up in <name>-jobs-manager, <name>-auto-upgrade, <name>-resource-monitor on a customer's production cluster, visible to anyone running kubectl there.

The change

A short docs/INSTALL.md section stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant means release == namespace still holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.

The claim is a machine check, not prose

The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so scripts/tests/release-name-equals-namespace.sh:

  • reads both arguments out of the installer's own invocation and asserts they are the same expression, whatever that expression is — so renaming the variable keeps it green, and passing a different value does not. It holds no copy of the expected name;
  • fails if the doc section is deleted, because a guard defending nothing is not a guard;
  • fails closed — an unreadable installer, a missing invocation, or an unparseable argument are findings, not agreement.

Wired into DRIFT_GUARDS (21 entries; the target counts its iterations and refuses to report green on fewer).

Worth recording: the first version of this guard matched a comment mentioning helm upgrade --install on line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.

Mutation evidence

mutationresult
installer passes a different release nameFAIL — prints both values
the doc drops the convention sectionFAIL — "protecting nothing"
the invocation disappears entirelyFAIL — refuses to report agreement
baseline, restoredPASS

make drift: all 21 guards green, with the new one verified to have executed. shellcheck clean.

Deliberately not included: fullnameOverride

I built it, proved the default render byte-identical, and reverted it. Filed as tracebloc/backend#2626 with the measurement:

  • 153.Release.Name sites across the templates, in five classes — only two of which may follow an override.
  • One of the others feeds RELEASE_NAME into helm status / helm rollback in the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failure backend#2620 is about.
  • Half-routed, the chart renders prod-auto-upgrade beside myrel-jobs-manager. A partial rename is worse than no rename.

It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.

Note

Trivially conflicts with #860 on the DRIFT_GUARDS line — both append a guard. Whichever merges second takes the other's entry.

Ticket: tracebloc/backend#2621


Note

Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.

Overview
Documents the Helm release name must match the namespace convention in docs/INSTALL.md (aligned with the bundled installer’s helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).

Adds scripts/tests/release-name-equals-namespace.sh as a drift guard: it parses the installer’s real helm upgrade --install block (line-anchored so comments don’t fake a match), asserts the release and --namespace arguments are the same expression, and fails if the doc section is removed. The guard is wired into DRIFT_GUARDS in the Makefile (21 guards).

Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.

…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@LukasWodkaLukasWodka self-assigned this Aug 27, 2026

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.

Comment threadscripts/tests/release-name-equals-namespace.sh Outdated
LukasWodkaand others added 2 commits August 27, 2026 09:06
The repo pipefail early-close check failed on the two parse lines in
scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an
assignment piping into head inherits head SIGPIPE kill once it closes early,
so a SUCCESSFUL parse could abort the guard - in a file whose whole point is
failing closed on an unparseable installer.
Capture-then-slice instead: collect every match with a here-string, then take
the first line with parameter expansion. No pipe, so no early-closing reader.
Still fails closed - no match leaves the variable empty and the existing -n
guards turn that into a finding.
Verified at CI severity, not narrower: org-github pipefail-early-close.awk
reports 0 findings across every .sh in the tree (was 2), and reintroducing
the pipe makes it report the line again. bash -n and shellcheck -S warning -x
clean. make check green. The guard still reddens on all four mutations -
different release name, changed --namespace, invocation removed, doc claim
deleted - and is green restored; every mutation anchor asserted applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@aptraceblocaptracebloc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.

Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.

Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).

— drafted with Claude Code

@LukasWodka
LukasWodka merged commit 1260229 into developAug 27, 2026
48 checks passed
@LukasWodka
LukasWodka deleted the feat/2621-release-name-convention branch August 27, 2026 08:13
aptracebloc added a commit that referenced this pull request Aug 27, 2026
…nstall (#863) (#871)
`Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested
anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not
registered." This is a harness race, not a chart defect. k3s registers its
bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the
resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes
report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready
nodes` and then helm-installed, so on a fast runner the install beat the addon
and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861
passed at 51s, neither touching the chart or these scripts).
Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and
call it after node-Ready, before the helm install, in every harness that installs
a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and
e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which
carries the preflight too). The helper POLLS for the APIService to EXIST first —
`kubectl wait` errors NotFound on a not-yet-created named object, so a bare
`kubectl wait --for=condition=Available` would just swap one red for another in
the same window — then best-effort waits for Available. It mirrors the production
installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553),
which faces the identical race. Rejected the weaker options (resourceMonitor:false
/ metricsServerPreflight:false): both go green only by deleting the #823 preflight
coverage this seal-check exists to exercise on a real cluster.
New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait
before their first install; the wait polls-for-existence before the condition
wait) so the guard cannot silently drift back out.
Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact
resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed
the helper blocks until registered+Available and the preflight then renders
satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical
to develop tip (zero failures added).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@LukasWodka@aptracebloc