Uh oh!
There was an error while loading. Please reload this page.
docs(install): write down the release-name convention, and guard it (backend#2621) - #861
Conversation
…backend#2621)
The installer has always named the release after the namespace —
`helm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"`, both
defaulting to `tracebloc` — and it explains why at the point of decision: the
client is identified to the backend by its clientId, "so we don't ask the user to
invent one". So the self-service path is consistent by construction.
It was written down nowhere. `docs/INSTALL.md`'s only example was
`helm install my-tracebloc …`, and `docs/migration-tools/tenant-config.example.env`
records the resulting mess as a fact to work around rather than a defect —
"The Helm release name (NOT always the namespace name; tenant-a's release is
`tracebloc`)". Hand-installed edges therefore diverged, and because HELM CANNOT
RENAME A RELEASE, every divergence is permanent: fifteen resource names are
prefixed with whatever was typed once, and correcting it means uninstall +
reinstall with the downtime and PV re-binding that implies.
The concrete case that prompted this: an engineer's given name ends up in
`<name>-jobs-manager`, `<name>-auto-upgrade`, `<name>-resource-monitor` on a
customer's production cluster, visible to anyone who runs kubectl there.
So this documents the convention, including the case the installer does not
cover — a multi-tenant cluster, where one namespace per tenant means
`release == namespace` still holds — and says plainly what not to do: not a
person's name, not a bare environment on a shared cluster, and keep it short
because Kubernetes truncates at 63 and the chart appends ~30 of component suffix.
AND THE CLAIM IS A MACHINE CHECK, not prose. The new section asserts something
about code ("the bundled installer already does this") in the document an
operator reads first — which is exactly the shape that decays into advice for
behaviour the code stopped having (backend#1729 rule 7).
`scripts/tests/release-name-equals-namespace.sh` reads BOTH arguments out of the
installer's own invocation and asserts they are the SAME EXPRESSION, whatever
that expression is — so renaming the variable keeps it green and passing a
different value does not. It holds no copy of the expected name. It also fails if
the doc section is deleted, because a guard defending nothing is not a guard.
Fails closed: an unreadable installer, a missing invocation, or an unparseable
argument are findings, not agreement.
Worth recording: the first version of this guard matched a COMMENT mentioning
`helm upgrade --install` on line 1025 rather than the invocation on line 2272 —
the same "prose satisfies a structural guard" class the file exists to prevent,
hit while writing it. It is now anchored to the start of a line, and the comment
says why.
Three mutations run, baseline restored green after each:
installer passes a different release name -> FAIL, prints both values
the doc drops the convention section -> FAIL, "protecting nothing"
the invocation disappears entirely -> FAIL, refuses to report agreement
make drift: all 21 guards green (verified the new one executed).
shellcheck clean.
NOT INCLUDED, deliberately: `fullnameOverride`. It is filed as
tracebloc/backend#2626 with the measurement — 153 `.Release.Name` sites across
five classes, of which only two may follow an override, and one of the others
feeds `RELEASE_NAME` into `helm status`/`helm rollback` in the auto-upgrade
script. A blanket substitution catches that env var on the first pass (mine did)
and breaks the auto-upgrade — the exact failure backend#2620 is about. A partial
rename is worse than none, so it needs a completeness guard rather than a helper.
Ticket: tracebloc/backend#2621
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 2f98d4a. Configure here.
Uh oh!
There was an error while loading. Please reload this page.
The repo pipefail early-close check failed on the two parse lines in scripts/tests/release-name-equals-namespace.sh. Under set -euo pipefail an assignment piping into head inherits head SIGPIPE kill once it closes early, so a SUCCESSFUL parse could abort the guard - in a file whose whole point is failing closed on an unparseable installer. Capture-then-slice instead: collect every match with a here-string, then take the first line with parameter expansion. No pipe, so no early-closing reader. Still fails closed - no match leaves the variable empty and the existing -n guards turn that into a finding. Verified at CI severity, not narrower: org-github pipefail-early-close.awk reports 0 findings across every .sh in the tree (was 2), and reintroducing the pipe makes it report the line again. bash -n and shellcheck -S warning -x clean. make check green. The guard still reddens on all four mutations - different release name, changed --namespace, invocation removed, doc claim deleted - and is green restored; every mutation anchor asserted applied. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts: # Makefile
aptracebloc
left a comment
There was a problem hiding this comment.
Approving. The "release name == namespace" convention is accurate (the installer defaults TB_NAMESPACE and passes one string for both the release arg and --namespace), and the new release-name-equals-namespace.sh drift-guard is genuinely derived-not-restated, fails closed, and is mutation-proof — I verified all three: release≠namespace fails, doc section deleted fails, real invocation removed fails. It's shellcheck-clean and the DRIFT_GUARDS count self-derives, so no hardcoded-count hazard. Green on make drift.
Two low docs-precision nits only (non-blocking): "Every resource … is prefixed" is slightly overbroad (resource-monitor RBAC uses the release as a suffix), and "Kubernetes truncates names at 63" is imprecise (the API server generally rejects >63; Helm caps release names at 53). Neither affects the operator takeaway.
Operational note for whoever merges after client#860: both touch the DRIFT_GUARDS line, so the second to merge must re-append the other's guard entry (and, given the self-count, confirm exp == ran).
— drafted with Claude Code
Uh oh!
There was an error while loading. Please reload this page.
…nstall (#863) (#871) `Seal-check egress-enforcement (k3d)` failed intermittently BEFORE it tested anything: "resourceMonitor is enabled but the metrics.k8s.io/v1beta1 API is not registered." This is a harness race, not a chart defect. k3s registers its bundled metrics-server addon — and the v1beta1.metrics.k8s.io APIService the resource-monitor preflight looks up (client#823) — ASYNCHRONOUSLY, after nodes report Ready. The e2e harnesses gated only on `kubectl wait --for=condition=Ready nodes` and then helm-installed, so on a fast runner the install beat the addon and the preflight `fail`ed the whole release (#862 false-failed at 26s while #861 passed at 51s, neither touching the chart or these scripts). Fix: add e2e_wait_for_metrics_apiservice to scripts/tests/lib/e2e-common.sh and call it after node-Ready, before the helm install, in every harness that installs a preflight-carrying chart directly: e2e-seal-check.sh, e2e-full-seal.sh, and e2e-auto-upgrade.sh (whose first install is the last PUBLISHED chart, which carries the preflight too). The helper POLLS for the APIService to EXIST first — `kubectl wait` errors NotFound on a not-yet-created named object, so a bare `kubectl wait --for=condition=Available` would just swap one red for another in the same window — then best-effort waits for Available. It mirrors the production installer's _wait_for_metrics_apiservice (lib/install-client-helm.sh, client#553), which faces the identical race. Rejected the weaker options (resourceMonitor:false / metricsServerPreflight:false): both go green only by deleting the #823 preflight coverage this seal-check exists to exercise on a real cluster. New e2e-metrics-apiservice-wait.bats pins the invariant (all three harnesses wait before their first install; the wait polls-for-existence before the condition wait) so the guard cannot silently drift back out. Verified on real k3d (rancher/k3s:v1.36.3-k3s1): reproduced the exact resource-monitor-daemonset.yaml:69 fail when the APIService is absent; confirmed the helper blocks until registered+Available and the preflight then renders satisfied-by-apiservice. `make lint` clean; full bats suite failure set identical to develop tip (zero failures added). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

Closes tracebloc/backend#2621
What was actually wrong
The installer has always named the release after the namespace:
…defaulting both to
tracebloc, and it explains itself at the point of decision (line 1862):So the self-service path is consistent by construction. The defect was that this was written down nowhere.
docs/INSTALL.md's only example washelm install my-tracebloc …, anddocs/migration-tools/tenant-config.example.envrecords the consequence as a fact to work around rather than a defect to fix:Hand-installed edges diverged accordingly — and because Helm cannot rename a release, every divergence is permanent. Fifteen resource names are prefixed with whatever was typed once; correcting it means uninstall + reinstall, with downtime and PV re-binding.
The case that prompted it: an engineer's given name ends up in
<name>-jobs-manager,<name>-auto-upgrade,<name>-resource-monitoron a customer's production cluster, visible to anyone runningkubectlthere.The change
A short
docs/INSTALL.mdsection stating the convention, including the case the installer doesn't cover — a multi-tenant cluster, where one namespace per tenant meansrelease == namespacestill holds — plus what not to do: not a person's name; not a bare environment on a shared cluster; keep it short, because Kubernetes truncates at 63 characters and the chart appends ~30 of component suffix.The claim is a machine check, not prose
The new section asserts something about code in the document an operator reads first. That is the shape that decays into advice for behaviour the code no longer has, so
scripts/tests/release-name-equals-namespace.sh:Wired into
DRIFT_GUARDS(21 entries; the target counts its iterations and refuses to report green on fewer).Worth recording: the first version of this guard matched a comment mentioning
helm upgrade --installon line 1025 rather than the real invocation on line 2272 — the same "prose satisfies a structural guard" class the file exists to prevent, hit while writing it. It is now anchored to the start of a line, and the comment says why.Mutation evidence
make drift: all 21 guards green, with the new one verified to have executed.shellcheckclean.Deliberately not included:
fullnameOverrideI built it, proved the default render byte-identical, and reverted it. Filed as
tracebloc/backend#2626with the measurement:.Release.Namesites across the templates, in five classes — only two of which may follow an override.RELEASE_NAMEintohelm status/helm rollbackin the auto-upgrade script. A blanket substitution catches that env var on the first pass — mine did — and breaks the auto-upgrade, which is the exact failurebackend#2620is about.prod-auto-upgradebesidemyrel-jobs-manager. A partial rename is worse than no rename.It needs a completeness guard, not a helper. That is a separate change with a separate risk profile.
Note
Trivially conflicts with #860 on the
DRIFT_GUARDSline — both append a guard. Whichever merges second takes the other's entry.Ticket: tracebloc/backend#2621
Note
Low Risk
Documentation and a read-only CI drift check; no chart or installer behavior changes.
Overview
Documents the Helm release name must match the namespace convention in
docs/INSTALL.md(aligned with the bundled installer’shelm upgrade --install "$TB_NAMESPACE" … --namespace "$TB_NAMESPACE"), with examples for default and multi-tenant installs and practical naming guidance (avoid personal names, avoid bare env names on shared clusters, stay DNS-1123-short).Adds
scripts/tests/release-name-equals-namespace.shas a drift guard: it parses the installer’s realhelm upgrade --installblock (line-anchored so comments don’t fake a match), asserts the release and--namespacearguments are the same expression, and fails if the doc section is removed. The guard is wired intoDRIFT_GUARDSin the Makefile (21 guards).Reviewed by Cursor Bugbot for commit 4576e58. Bugbot is set up for automated code reviews on this repo. Configure here.