Skip to content

Develop - #2

Merged
saadqbal merged 2 commits into
mainfrom
develop
Aug 28, 2025
Merged

Develop#2
saadqbal merged 2 commits into
mainfrom
develop

Conversation

@saadqbal

Copy link
Copy Markdown
Contributor

No description provided.

@saadqbal
saadqbal merged commit 2bdf4fa into mainAug 28, 2025
saadqbal added a commit that referenced this pull request Apr 28, 2026
…ot follow-up #2 (#80)
* fix(client): release-scope SCC SA refs (v1.2.2)
Bugbot caught a High-severity miss in v1.2.0's release-scoping work
(PR #72). The OpenShift SCC template was the one resource-monitor file
not updated when the literal `tracebloc-resource-monitor` ServiceAccount
name moved to `<Release.Name>-resource-monitor`. On OpenShift the SCC
granted access to a SA name that no longer existed, so the resource-
monitor DaemonSet pods would fail to launch (no SCC -> can't mount
hostPath /proc and /sys for node metrics).
The SCC's metadata.name + ClusterRole.name + ClusterRoleBinding.name
were ALREADY release-scoped (`tracebloc-resource-monitor-<release>` /
`tracebloc-resource-monitor-scc-<release>`), so this slipped through —
casual reading suggested it was already done.
Touchpoints in resource-monitor-scc.yaml:
- users[0]: now {{ include "tracebloc.resourceMonitorName" . }}
- ClusterRoleBinding subjects[0].name: same helper
- All `app: tracebloc-resource-monitor` labels: same helper, for
consistency with the rest of the chart's resource-monitor templates
- Updated the kubernetes.io/description SCC annotation prose so the
literal name doesn't appear there either (cosmetic, but easier to
audit "no literal references" with a single grep).
Tests:
- platform_test.yaml gains 3 new cases: SCC users[0] points at
release-scoped SA, ClusterRoleBinding subject does too, and two
releases (stg + cisco/hasan-prod) produce non-colliding SA references.
- node_agents_namespace_test.yaml had a regression assertion checking
the OLD literal name in users[0]; updated to the new release-scoped
form (`RELEASE-NAME-resource-monitor`, helm-unittest's default
release name when none is set).
- 98 -> 102 passing.
Verified end-to-end with two side-by-side `helm template` runs:
- stg -> users[0] = system:serviceaccount:tracebloc-node-agents:stg-resource-monitor
- hasan-prod -> users[0] = system:serviceaccount:tracebloc-node-agents:hasan-prod-resource-monitor
Chart bumped 1.2.1 -> 1.2.2 (patch — restores OpenShift parity that
v1.2.0 inadvertently broke).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: NOTES.txt rename + generator chart-version drift (v1.2.3)
Bugbot follow-up to the v1.2.0/1.2.2 rename work. Two fresh issues:
1. (Medium) NOTES.txt:9 still hardcoded the literal
`tracebloc-resource-monitor` for the resource-monitor DaemonSet
display, while the actual DaemonSet name has been
`<release>-resource-monitor` since v1.2.0. Operators see one name
in the post-install banner and a different name when they
`kubectl get ds`. Now routes through the same
tracebloc.resourceMonitorName helper as the rest of the chart.
2. (Low) docs/migration-tools/generate.sh hardcoded
`app.kubernetes.io/version: "1.1.0"` and `helm.sh/chart: client-1.1.0`
on every pre-create PVC. The chart has moved through 1.1.0 → 1.2.3,
and operators running generate.sh today get PVC labels stuck at
1.1.0 even though the install ahead is 1.2.3. Helm adoption itself
is unaffected (it keys on meta.helm.sh/release-name, not the chart
label), but the labels lie until a subsequent upgrade reconciles
them, and `kubectl get pvc -L helm.sh/chart` is misleading during
migration debugging. Fixed by reading name + version from
client/Chart.yaml at generate time.
Plus a few stale prose references caught while auditing the same path
(no functional impact, but the doc was directing operators at "client
fix in 1.2.0" as if it were still pending):
- generate.sh inline comment on `resourceMonitor: false` rephrased
from "until client-1.2.0 is published" to "until you have verified
the chart you're installing is 1.2.0+"
- migrate-tenant.sh banner relabelled from "v1.1.0 spec sanity" to
"mysql spec sanity (v1.1.0+ shape: ...)"
- README.md skip table cell on `resourceMonitor: false` rewritten to
reflect that 1.2.0+ has shipped — operators on >=1.2.0 can flip it
to true without colliding with the stg release
Tests: 102 → 105 passing. New `client/tests/notes_test.yaml` covers:
- Release-scoped resource-monitor name appears in NOTES.txt
- A different release renders a different name (proves the helper
isn't accidentally hardcoded)
- Negative regex guards against the literal `tracebloc-resource-monitor`
reappearing followed by a non-suffix character (i.e. the bare
pre-1.2.3 form, while still letting the SCC line `tracebloc-
resource-monitor-<release>` further down the file pass)
- `resourceMonitor: false` removes the line entirely
End-to-end smoke of generate.sh confirms PVCs ship with the live chart
version (`helm.sh/chart: client-1.2.3` after this commit, verified
against /tmp/tracebloc-migration-<demo>/pvcs.yaml).
Stacked on PR #78 (v1.2.2 SCC fix), so this branch already contains
the SCC SA-ref rename. Once #78 lands the diff against develop will
reduce to just this commit.
Chart bumped 1.2.2 → 1.2.3 (patch — operator-facing string fix +
tooling correctness).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
saadqbal added a commit that referenced this pull request Apr 28, 2026
* Merge pull request #71 from tracebloc/docs/migrations-correct-option-b
docs(migrations): correct Option B + add hasan-prod case + active-jobs pre-flight
* chore: add default CODEOWNERS for auto-reviewer assignment (#73)
* ci: add kanban closure-routing caller workflow (#75)
* fix(client): release-scope resource-monitor names so multiple releases coexist (v1.2.0) (#72)
Two client releases on the same cluster could not both deploy the
resource-monitor DaemonSet because several resources templated into the
shared tracebloc-node-agents namespace used the literal name
`tracebloc-resource-monitor` rather than a release-scoped name. The
second `helm install` failed with:
Error: ServiceAccount "tracebloc-resource-monitor" in namespace
"tracebloc-node-agents" exists and cannot be imported into the current
release: invalid ownership metadata; ... must equal "hasan-prod":
current value is "stg".
Surfaced during the 2026-04-27 hasan-prod migration on
tracebloc-templates-prod; worked around at the time by setting
resourceMonitor: false on the second release, which means prod customers
currently lose their per-CLIENT_ID metric stream until this lands.
What changed:
- New helper `tracebloc.resourceMonitorName` -> `<Release.Name>-resource-monitor`,
centralised in _helpers.tpl alongside the existing per-release name
helpers (secretName, serviceAccountName, etc.).
- DaemonSet metadata.name, spec.selector.matchLabels.app, pod label
app=, and spec.template.spec.serviceAccountName all now go through
the helper. The selector + pod label have to move together because
DaemonSet selectors are namespace-scoped: two DaemonSets in
tracebloc-node-agents both selecting `app: tracebloc-resource-monitor`
would each grab the other's pods, which is worse than the surface bug.
- ServiceAccount metadata.name (resource-monitor-rbac.yaml) goes through
the helper. ClusterRole / ClusterRoleBinding / Role / RoleBinding
metadata.name were already release-scoped (`tracebloc-resource-monitor-<release>`)
and stay as-is to avoid an unnecessary ClusterRole rename for upgrading
installs. Only the *subject* names in (Cluster)RoleBinding change to
point at the new SA.
- Mirrored secrets (CLIENT_ID + dockerconfigjson) in tracebloc-node-agents:
the secret names were already release-scoped via
tracebloc.secretName / tracebloc.registrySecretName so they did not
collide. Their `app` label was the literal value, which is harmless on
uniquely-named resources but inconsistent — updated for consistency.
- Chart bumped 1.1.0 -> 1.2.0. Per-release naming of cluster-singleton
resources is a behaviour change for existing installs (DaemonSet name,
ServiceAccount name, and selector label all change), so a minor bump
signals that operators should review.
Tests: 93 -> 98. New cases cover:
- DaemonSet name + selector + serviceAccountName all release-scoped
- ServiceAccount name release-scoped
- ClusterRoleBinding subject points at the release-scoped SA
- A second `helm template` with a different release name produces
non-colliding names
Verified end-to-end via `helm template stg ./client` and
`helm template hasan-prod ./client` on the same chart: ServiceAccount,
DaemonSet, and ClusterRoleBinding subject names all diverge per release.
Upgrade path from 1.1.0:
The DaemonSet and ServiceAccount rename triggers a Helm three-way merge
that DELETEs the old `tracebloc-resource-monitor` resource and CREATEs
the new release-scoped one. ~30-60s gap on each node where resource
metrics are not collected. DaemonSet selector is immutable, so the
delete-then-create path is what we want — helm upgrade handles this
automatically because the names diverge in the stored manifest. No
manual orphan cleanup needed.
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* fix(client): allow training pods to reach mysql-client (v1.2.1) (#76)
The training-egress NetworkPolicy added in v1.1.0 only permitted DNS and
external TCP/443. Training pods load their dataset from the in-namespace
mysql-client over TCP/3306 (core/utils/database.py::load_dataframe_from_sql_table),
so under any CNI that actually enforces NetworkPolicy the connect failed
with errno 111 and the Job CrashLoopBackOff'd before the first batch:
Database connection failed: 2003 (HY000): Can't connect to MySQL server
on 'mysql-client:3306' (111)
RuntimeError: Database connection is not available for load_dataframe_from_sql_table
Surfaced on a fresh client install (k3d / k3s, which enforces policy via
the built-in kube-router) where jobs-manager could reach mysql but every
training Job spawned with tracebloc.io/workload=training could not.
Add a third egress rule scoped to podSelector {app: mysql-client} on
TCP/3306. Same-namespace by default (no namespaceSelector), so it stays
tight to the chart's own mysql pod and does not open the namespace
generally. The egress[1] /32 ipBlock comment is updated to note that
MySQL is now explicitly re-permitted by egress[2].
Verified on a k3d cluster: pre-fix nc to mysql-client:3306 from a pod
with the training label was refused; post-fix it connects.
* docs(migration-tools): tenant migration runbook for eks-1.0.x → client-1.x (#74)
* docs(migration-tools): tenant migration runbook for eks-1.0.x -> client-1.x
Captures the operational tooling validated during the 2026-04-27 stg and
hasan-prod migrations and generalises it for the remaining tenants
(bmw, cisco, charite) and any future tenant on the legacy chart family.
What's here:
- README.md walks the workflow + recommended ordering for the pending
set + skip rationale for chart toggles (resourceMonitor: false,
priorityClass.create: false, etc).
- generate.sh consumes a tenant-config.env (gitignored) and emits, per
tenant, /tmp/tracebloc-migration-<tenant>/{values,storageclass,pvcs}.yaml.
Refuses to expand placeholder __FOO__ rows so an operator running
generate.sh against the unmodified template fails fast.
- migrate-tenant.sh is the parameterised runbook. `phase1` is
non-destructive (mysqldump-then-chunked-cp, AWS Backup on-demand
recovery point, dry-run render). `phase2` is one-shot per tenant
(helm uninstall, claimRef clear, SC re-create, PVC pre-create with
release-scoped Helm ownership stamp, helm install, verify mysql data
+ keep annotation in stored manifest).
- tenant-config.example.env is the template; populated copy is the
secret-bearing artifact and must stay local.
No real secrets in any committed file:
- DOCKER_PASSWORD placeholder (__DOCKER_HUB_PERSONAL_ACCESS_TOKEN__)
- per-tenant CLIENT_ID / CLIENT_PASSWORD placeholders
- MYSQL_ROOT_PW placeholder (it's image-baked; required from env at
runtime, no committed default)
- .gitignore now excludes docs/migration-tools/tenant-config.env
(only the .example variant is tracked)
Operational notes:
- Every kubectl/helm call passes --context explicitly. The 2026-04-27
prod run hit a context-drift bug mid-migration; the explicit form
is a hard requirement.
- values.yaml ships with resourceMonitor: false. Flip true after the
release-scoped resource-monitor names land in client-1.2.0 (separate
PR). Until then the shared SA in tracebloc-node-agents collides with
the stg release.
- Phase 1 is idempotent and re-runnable. Phase 2 is destructive and
one-shot per tenant. Operators should pause and eyeball Phase 1
outputs before running Phase 2 — that's deliberately not automated.
Once all four pending tenants are on client-1.x, this directory is
historical. client-1.x -> client-1.y upgrades follow plain `helm upgrade`
because the new chart already templates `helm.sh/resource-policy: keep`
on PVCs, so the migration protocol isn't needed for routine upgrades.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(migration-tools): address bugbot review feedback on PR #74
Three issues flagged by Cursor Bugbot on the migration scripts:
* migrate-tenant.sh used macOS-only `md5 -q` and `stat -f%z` for chunked-cp
verification (HIGH). Linux operators would abort Phase 1 mid-transfer.
Add portable `_md5` and `_size` helpers that pick md5sum on Linux,
fall back to md5(1) on macOS, and use `wc -c` instead of stat for size.
* generate.sh placeholder gate inspected only CLIENT_ID + CLIENT_PASSWORD
+ PV_MYSQL, missing PV_LOGS, PV_DATA, SC_NAME, and DOCKER_PASSWORD
(MEDIUM). Literal `__FOO__` placeholders silently rendered into
values.yaml/pvcs.yaml and only blew up at kubectl apply / helm install
time. Iterate over every per-row field, plus a one-shot global check
for DOCKER_PASSWORD before the loop. Error messages now name the
offending field.
* Phase 2.5 readiness loop was an unbounded `while :; do … sleep 5; done`
(MEDIUM). After the destructive helm uninstall, a non-converging
install (image-pull error, mysql kill-loop recurrence, missing PVC
binding) hung the script forever instead of surfacing the failure.
Add a wall-clock deadline — default 600s, override via READY_TIMEOUT —
and exit 1 with the last-seen pod state on timeout.
* fix(migration-tools): address bugbot follow-up on PR #74
Two more issues raised on the previous fix commit:
* Readiness wait loop aborted on empty pod list (HIGH). With `set -euo
pipefail`, the routine post-install window where no pods are visible
yet caused `grep -c .` to exit 1, killing the script on the very first
iteration before the wall-clock deadline could ever fire — defeating
the bounded-wait intent. Guard the empty case explicitly. `wc -l`
alone is also wrong because `echo ""` prints a newline.
* MYSQL_ROOT_PW skipped the placeholder check that DOCKER_PASSWORD,
CLIENT_*, and PV_* now have (LOW). An operator who copied the example
without editing this row passed the non-empty gate, then the literal
__LEGACY_MYSQL_ROOT_PW__ went into mysqldump and Phase 1 blew up
partway through with an opaque "Access denied" inside kubectl exec.
Add the same `*__*__*` case guard right after the non-empty check.
* fix(migration-tools): make EFS_FS_OVERRIDE actually override (PR #74)
The pre-source assignment
EFS_FS="${EFS_FS_OVERRIDE:-fs-06b3faf51675ff9f9}"
was a no-op: `source "$CONFIG"` runs immediately after and the example
config (and any real tenant-config.env derived from it) unconditionally
sets EFS_FS=fs-06b3faf51675ff9f9, so the env override was clobbered every
time. Operators thinking they were targeting a non-default EFS would
silently start AWS Backup on-demand jobs against the hard-coded prod
filesystem.
Move the override knob to AFTER source where env genuinely wins, drop
the hard-coded fallback, and require EFS_FS to be set somewhere (config
or override) before continuing.
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* fix(client): release-scope SCC SA refs (v1.2.2) (#78)
Bugbot caught a High-severity miss in v1.2.0's release-scoping work
(PR #72). The OpenShift SCC template was the one resource-monitor file
not updated when the literal `tracebloc-resource-monitor` ServiceAccount
name moved to `<Release.Name>-resource-monitor`. On OpenShift the SCC
granted access to a SA name that no longer existed, so the resource-
monitor DaemonSet pods would fail to launch (no SCC -> can't mount
hostPath /proc and /sys for node metrics).
The SCC's metadata.name + ClusterRole.name + ClusterRoleBinding.name
were ALREADY release-scoped (`tracebloc-resource-monitor-<release>` /
`tracebloc-resource-monitor-scc-<release>`), so this slipped through —
casual reading suggested it was already done.
Touchpoints in resource-monitor-scc.yaml:
- users[0]: now {{ include "tracebloc.resourceMonitorName" . }}
- ClusterRoleBinding subjects[0].name: same helper
- All `app: tracebloc-resource-monitor` labels: same helper, for
consistency with the rest of the chart's resource-monitor templates
- Updated the kubernetes.io/description SCC annotation prose so the
literal name doesn't appear there either (cosmetic, but easier to
audit "no literal references" with a single grep).
Tests:
- platform_test.yaml gains 3 new cases: SCC users[0] points at
release-scoped SA, ClusterRoleBinding subject does too, and two
releases (stg + cisco/hasan-prod) produce non-colliding SA references.
- node_agents_namespace_test.yaml had a regression assertion checking
the OLD literal name in users[0]; updated to the new release-scoped
form (`RELEASE-NAME-resource-monitor`, helm-unittest's default
release name when none is set).
- 98 -> 102 passing.
Verified end-to-end with two side-by-side `helm template` runs:
- stg -> users[0] = system:serviceaccount:tracebloc-node-agents:stg-resource-monitor
- hasan-prod -> users[0] = system:serviceaccount:tracebloc-node-agents:hasan-prod-resource-monitor
Chart bumped 1.2.1 -> 1.2.2 (patch — restores OpenShift parity that
v1.2.0 inadvertently broke).
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* fix: NOTES.txt rename + generator chart-version drift (v1.2.3) — bugbot follow-up #2 (#80)
* fix(client): release-scope SCC SA refs (v1.2.2)
Bugbot caught a High-severity miss in v1.2.0's release-scoping work
(PR #72). The OpenShift SCC template was the one resource-monitor file
not updated when the literal `tracebloc-resource-monitor` ServiceAccount
name moved to `<Release.Name>-resource-monitor`. On OpenShift the SCC
granted access to a SA name that no longer existed, so the resource-
monitor DaemonSet pods would fail to launch (no SCC -> can't mount
hostPath /proc and /sys for node metrics).
The SCC's metadata.name + ClusterRole.name + ClusterRoleBinding.name
were ALREADY release-scoped (`tracebloc-resource-monitor-<release>` /
`tracebloc-resource-monitor-scc-<release>`), so this slipped through —
casual reading suggested it was already done.
Touchpoints in resource-monitor-scc.yaml:
- users[0]: now {{ include "tracebloc.resourceMonitorName" . }}
- ClusterRoleBinding subjects[0].name: same helper
- All `app: tracebloc-resource-monitor` labels: same helper, for
consistency with the rest of the chart's resource-monitor templates
- Updated the kubernetes.io/description SCC annotation prose so the
literal name doesn't appear there either (cosmetic, but easier to
audit "no literal references" with a single grep).
Tests:
- platform_test.yaml gains 3 new cases: SCC users[0] points at
release-scoped SA, ClusterRoleBinding subject does too, and two
releases (stg + cisco/hasan-prod) produce non-colliding SA references.
- node_agents_namespace_test.yaml had a regression assertion checking
the OLD literal name in users[0]; updated to the new release-scoped
form (`RELEASE-NAME-resource-monitor`, helm-unittest's default
release name when none is set).
- 98 -> 102 passing.
Verified end-to-end with two side-by-side `helm template` runs:
- stg -> users[0] = system:serviceaccount:tracebloc-node-agents:stg-resource-monitor
- hasan-prod -> users[0] = system:serviceaccount:tracebloc-node-agents:hasan-prod-resource-monitor
Chart bumped 1.2.1 -> 1.2.2 (patch — restores OpenShift parity that
v1.2.0 inadvertently broke).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix: NOTES.txt rename + generator chart-version drift (v1.2.3)
Bugbot follow-up to the v1.2.0/1.2.2 rename work. Two fresh issues:
1. (Medium) NOTES.txt:9 still hardcoded the literal
`tracebloc-resource-monitor` for the resource-monitor DaemonSet
display, while the actual DaemonSet name has been
`<release>-resource-monitor` since v1.2.0. Operators see one name
in the post-install banner and a different name when they
`kubectl get ds`. Now routes through the same
tracebloc.resourceMonitorName helper as the rest of the chart.
2. (Low) docs/migration-tools/generate.sh hardcoded
`app.kubernetes.io/version: "1.1.0"` and `helm.sh/chart: client-1.1.0`
on every pre-create PVC. The chart has moved through 1.1.0 → 1.2.3,
and operators running generate.sh today get PVC labels stuck at
1.1.0 even though the install ahead is 1.2.3. Helm adoption itself
is unaffected (it keys on meta.helm.sh/release-name, not the chart
label), but the labels lie until a subsequent upgrade reconciles
them, and `kubectl get pvc -L helm.sh/chart` is misleading during
migration debugging. Fixed by reading name + version from
client/Chart.yaml at generate time.
Plus a few stale prose references caught while auditing the same path
(no functional impact, but the doc was directing operators at "client
fix in 1.2.0" as if it were still pending):
- generate.sh inline comment on `resourceMonitor: false` rephrased
from "until client-1.2.0 is published" to "until you have verified
the chart you're installing is 1.2.0+"
- migrate-tenant.sh banner relabelled from "v1.1.0 spec sanity" to
"mysql spec sanity (v1.1.0+ shape: ...)"
- README.md skip table cell on `resourceMonitor: false` rewritten to
reflect that 1.2.0+ has shipped — operators on >=1.2.0 can flip it
to true without colliding with the stg release
Tests: 102 → 105 passing. New `client/tests/notes_test.yaml` covers:
- Release-scoped resource-monitor name appears in NOTES.txt
- A different release renders a different name (proves the helper
isn't accidentally hardcoded)
- Negative regex guards against the literal `tracebloc-resource-monitor`
reappearing followed by a non-suffix character (i.e. the bare
pre-1.2.3 form, while still letting the SCC line `tracebloc-
resource-monitor-<release>` further down the file pass)
- `resourceMonitor: false` removes the line entirely
End-to-end smoke of generate.sh confirms PVCs ship with the live chart
version (`helm.sh/chart: client-1.2.3` after this commit, verified
against /tmp/tracebloc-migration-<demo>/pvcs.yaml).
Stacked on PR #78 (v1.2.2 SCC fix), so this branch already contains
the SCC SA-ref rename. Once #78 lands the diff against develop will
reduce to just this commit.
Chart bumped 1.2.2 → 1.2.3 (patch — operator-facing string fix +
tooling correctness).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* docs(claude): require @saadqbal as PR assignee (#79)
Convention captured after a session-end ask. Every PR Claude opens for
this repo must be assigned to saadqbal — orphaned PRs without an
assignee fall through the review queue.
Pass --assignee @me on `gh pr create` (or --assignee saadqbal if running
unauthenticated). No exceptions.
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: lukasWuttke <54042461+LukasWodka@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
saadqbal added a commit that referenced this pull request Jun 4, 2026
…can't resolve (#191)
image-refresh silently skipped every tick when get_latest_digest returned
empty for the ghcr.io ingestor image (egress/proxy/firewall to ghcr.io, or a
blocked token endpoint) — never reaching the registry-drift branch that sets
the new digest. jobs-manager + pods-monitor pull from docker.io and refreshed
fine, so the CronJob looked healthy while the ingestor digest stayed pinned on
the install-time baseline. That's why the berlin-team arm64 install sat on the
amd64-only v0.3.1 digest even after :0.3 went multi-arch (#186 follow-up #2).
Now count consecutive ingestor-resolve failures on a deployment annotation:
- below imageRefresh.ingestorResolveFailureThreshold (default 3, ~45 min at the
15-min schedule) -> WARN + skip, as before (tolerate transient blips);
- at/above it -> ERROR with actionable guidance, a
tracebloc.io/ingestor-refresh-last-error annotation, and a non-zero exit so
the Job fails visibly in `kubectl get cronjob` / monitoring — the same
surfacing idiom Pass 2's stuck-rollout check already relies on;
- a successful resolve clears the streak.
Threshold is nil-guarded (default 3) for --reuse-values upgrades and
schema-validated (integer >= 1). The digest-resolution logic itself is
unchanged (verified correct: it returns the multi-arch index digest).
helm unittest 146/146, helm lint clean, shellcheck + sh -n clean.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
saadqbal added a commit that referenced this pull request Jun 4, 2026
* fix(resource-monitor): always grant read-only ClusterRole (decouple from clusterScope)
Under clusterScope: false the chart rendered only a namespace-scoped Role in
the release namespace. But the resource-monitor's code:
* calls core_v1_api.list_pod_for_all_namespaces(field_selector=spec.nodeName=...)
-- a CLUSTER-SCOPED list verb a namespaced Role can never satisfy; and
* read_namespaced_pod()s its OWN pod, which lives in
.Values.nodeAgents.namespace.name (NOT .Release.Namespace).
So with clusterScope: false the DaemonSet 403'd on startup and crashlooped
(70+ restarts observed on a live cluster). Per-node monitoring is intrinsically
cluster-scoped.
Always render the read-only ClusterRole + ClusterRoleBinding regardless of
clusterScope (get/list/watch on pods/nodes/namespaces + metrics; no write,
exec, or secret access). resourceMonitor: false still fully disables the
component. clusterScope continues to gate the training/jobs isolation footprint
elsewhere -- it must not leave the node monitor without permissions it cannot
run without.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(resource-monitor): assert always-cluster-scoped RBAC under clusterScope=false
Follow-up to the RBAC fix: node_agents_namespace_test.yaml still asserted the
old behavior (namespaced Role + RoleBinding in the release namespace when
clusterScope=false). Update that case to assert the corrected contract -- a
ClusterRole + ClusterRoleBinding always render (with no metadata.namespace),
while the subject SA still lives in the node-agents namespace.
The clusterScope=false path stays under test; only the asserted behavior
changes to match the fix. Verified with `helm unittest` (all resource-monitor
suites pass).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(rbac): grant `get` on configmaps/secrets to jobs-manager SA
The ingestion endpoint's orphan-resource verify path (client-runtime#52)
and missing-row self-heal (client-runtime#54) read the existing
ConfigMap/Secret on a create-409 to confirm content matches before
reuse. The Role/ClusterRole only granted `create`, so those reads
returned Forbidden and the endpoint 500'd instead of the intended
409/200-replay — verified live on the dev cluster.
Add `get` alongside `create` in both the ClusterRole (clusterScope:
true) and namespace Role (clusterScope: false) branches.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* ci(helm): guard that the pinned ingestor digest is multi-arch (closes#186) (#187)
Add a helm-ci job (ingestor-multiarch) that parses images.ingestor.digest and
fails the build unless it's a multi-arch index (linux/amd64 + linux/arm64).
Greenfield installs spawn the ingestor Job from this PINNED digest before
image-refresh first ticks, so an amd64-only pin breaks data ingestion on arm64
hosts (Apple Silicon, Graviton) with "no match for platform" / ImagePullBackOff.
This would have caught #160 (the amd64-only v0.3.1 pin) before it shipped.
ghcr.io/tracebloc/ingestor is public -> no secrets. Verified: passes on the
current multi-arch baseline (sha256:d361fa77, v0.3.2 / #184), fails on the old
amd64-only sha256:a0861ea9.
Note: the digest is already multi-arch on develop as of v0.3.2 (#184 — the same
d361fa77 index this PR previously bumped to), so #187 no longer touches
values.yaml; it adds only the regression guard so an amd64-only pin can't slip
back in.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#190): fail image-refresh loudly when the ingestor (ghcr) digest can't resolve (#191)
image-refresh silently skipped every tick when get_latest_digest returned
empty for the ghcr.io ingestor image (egress/proxy/firewall to ghcr.io, or a
blocked token endpoint) — never reaching the registry-drift branch that sets
the new digest. jobs-manager + pods-monitor pull from docker.io and refreshed
fine, so the CronJob looked healthy while the ingestor digest stayed pinned on
the install-time baseline. That's why the berlin-team arm64 install sat on the
amd64-only v0.3.1 digest even after :0.3 went multi-arch (#186 follow-up #2).
Now count consecutive ingestor-resolve failures on a deployment annotation:
- below imageRefresh.ingestorResolveFailureThreshold (default 3, ~45 min at the
15-min schedule) -> WARN + skip, as before (tolerate transient blips);
- at/above it -> ERROR with actionable guidance, a
tracebloc.io/ingestor-refresh-last-error annotation, and a non-zero exit so
the Job fails visibly in `kubectl get cronjob` / monitoring — the same
surfacing idiom Pass 2's stuck-rollout check already relies on;
- a successful resolve clears the streak.
Threshold is nil-guarded (default 3) for --reuse-values upgrades and
schema-validated (integer >= 1). The digest-resolution logic itself is
unchanged (verified correct: it returns the multi-arch index digest).
helm unittest 146/146, helm lint clean, shellcheck + sh -n clean.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* test(requests-proxy): add helm-unittest coverage for requests-proxy Deployment (#194)
requests-proxy-deployment.yaml was the only data-plane workload template
without a unit test. This suite pins the properties most costly to regress:
- security-context invariants (no SA-token automount, runAsNonRoot,
seccomp RuntimeDefault, runAsUser 1001, no privilege escalation,
drop ALL caps, read-only root filesystem) — see docs/SECURITY.md
- the single-replica / single gunicorn worker constraint (the pod token
registry is process-local; >1 worker silently shards token lookups)
- the docker.io/tracebloc/jobs-manager image source and port 8888
- the nil-guarded resource defaults, plus an override case that exercises
the default-through-dict fallthrough (guards the historic
`readOnlyRootFilesystem: trueresources:` newline-eating regression)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#196): allow training-pod egress to the requests-proxy (8888) (#197)
The training-egress NetworkPolicy denies all pod-to-pod / ClusterIP
egress (rule 2 excepts the cluster CIDRs) and re-permits only MySQL
(rule 3). When the requests-proxy architecture shipped — training pods
POST epoch results / FLOPs to requests-proxy-service:8888 instead of
holding Service Bus credentials — this template was never updated to
re-permit egress to the proxy. Result on every install with the policy
enabled: pods hit "requests-proxy-service:8888 ... [Errno 111]
Connection refused" at the first epoch finalize → CrashLoopBackOff →
all experiments fail.
Add rule 4 mirroring the MySQL rule: TCP/8888 to podSelector
app=requests-proxy (same namespace). Service selector + port from
templates/requests-proxy-service.yaml.
Verified: `helm template -f ci/bm-values.yaml --show-only
templates/network-policy-training.yaml` renders the new rule as valid
YAML.
Found live on a fresh client (tracebloc-amazon / k3d): jobs-manager
reached the proxy (HTTP 401) while training pods got connection-refused
— the only differentiator was this egress policy. Interim: live-patched
the cluster + suspended its auto-upgrade CronJob (so reuse-values
wouldn't revert the patch); re-enable once this lands + releases.
Closes#196.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Installer UX: drop PriorityClass, fix namespace, one-per-machine guard, surface version (#192)
* fix(chart): drop the data-plane PriorityClass by default
The cluster-scoped, fixed-name `tracebloc-data-plane` PriorityClass was the
only thing forcing one tracebloc client per cluster (a second release collided
on it with a cryptic Helm error) and blocking multiple tracebloc namespaces in
one BYO cluster. mysql doesn't need it: memory requests==limits (last evicted
under memory pressure), data on a PVC (eviction = transient restart, not data
loss), and a PDB guards voluntary disruptions. Its only unique benefit was
letting the scheduler preempt training jobs to keep mysql scheduled on a packed
node — a narrow case.
Default priorityClass.create=false + name="" so new installs template no
PriorityClass and mysql carries no priorityClassName. Opt back in
(create:true + name) on contended clusters, or reference an out-of-band one
(create:false + name:<existing>). helm-unittest updated; 144/144 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(installer): fixed namespace + one-per-machine guard; drop workspace prompt
The "Choose a workspace name" prompt asked the user to invent a label that
isn't their identity (the backend identifies a client by its credentials, not
this string — it's just the local k8s namespace / Helm release name; the
installer even discards the auth response body). It defaulted to a meaningless
"default" and was the field that collided on a second install.
- Drop the prompt; TB_NAMESPACE defaults to a fixed "tracebloc"
(env-overridable for advanced/GitOps setups).
- One-client-per-machine guard: after credentials verify, compare the entered
Client ID against any client already installed here (helm get values). Same
ID = a normal re-run/upgrade; a DIFFERENT ID hard-blocks with an explanation
and options (repair / switch via `k3d cluster delete` / use another machine)
instead of silently re-pointing the machine. This replaces the accidental
PriorityClass collision (now dropped) with an intentional, explained guard.
- Document the TB_NAMESPACE override; update bats (input sequences + 2 new
guard tests). bats 26/27 — the 1 failure is a pre-existing macOS-bash-3.2
quirk in _extract_yaml_value, unrelated (CI bash 5 passes it).
NOTE: install-k8s.ps1 + its Pester tests still need the same mirror (follow-up).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(installer-ps): mirror fixed namespace + one-per-machine guard (PowerShell)
Mirrors the bash change in install-k8s.ps1:
- drop the "Choose a workspace name" prompt; TB_NAMESPACE defaults to a fixed
"tracebloc" (override via $env:TB_NAMESPACE).
- one-client-per-machine guard: after credentials verify, compare the entered
Client ID against any client already installed here (helm get values); a
different ID hard-blocks with the same explanation/options as bash.
- Pester: 2 new guard tests (block-different / allow-same). The existing
Install-ClientHelm tests use dispatch-by-prompt Read-Host mocks, so the
prompt removal doesn't disturb them.
No pwsh locally -> verified via CI (Pester ubuntu+windows + PSScriptAnalyzer).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): scan all namespaces in the one-per-machine guard
The guard checked only the `tracebloc` namespace, so a client installed by an
older installer version (default namespace `default`, or a custom name) wasn't
detected -- a re-run could create a second coexisting client. Now enumerate all
client-chart releases (helm list -A) and compare each one's clientId, covering
both fresh and migrated installs. bash uses jq (already a dependency; falls
back to the tracebloc namespace if absent); PowerShell uses ConvertFrom-Json.
The block message names the namespace. bats + Pester guard tests updated.
Verified: bats green locally (jq path); PowerShell via CI Pester.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(installer): show client (chart) version in summary + --diagnose
Users had no easy way to see which client version they're on (the CLI isn't
shipped yet; `helm list` needs the namespace, and nothing surfaced it). Show
the chart version where they already look:
- install summary: a "Version" line next to Workspace.
- --diagnose: as the first console line + recorded in the bundle header
(the #1 thing support needs).
Adds a best-effort `_chart_version` / Get-ChartVersion helper (greps helm's
CHART column -> no jq). bash + PowerShell; bats + Pester coverage added.
Verified: summary.bats + diagnose.bats green locally; ps1 via CI Pester.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* chore: bump chart 1.4.4 → 1.4.5 to ship the training-egress proxy fix (#198)
1.4.4 is already published on the tracebloc.github.io/client Pages
channel and is what clusters run. The training-egress NetworkPolicy fix
(#197, allow training → requests-proxy:8888) merged to develop without a
version bump, so it is currently undeliverable: chart-releaser won't
overwrite the existing 1.4.4 release, and clusters already on 1.4.4 would
see no version change and pull nothing.
Bump to 1.4.5 (lockstep version/appVersion, matching 1.4.3/1.4.4 history)
so a v1.4.5 release publishes a new version that auto-upgrade actually
pulls. Chart-only change; no image change.
Ref #196 / #197.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: shujaat hasan <shujaathasan@shujaats-MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: shujaat_tracebloc <153823837+shujaatTracebloc@users.noreply.github.com>
Co-authored-by: lukasWuttke <54042461+LukasWodka@users.noreply.github.com>
saadqbal added a commit that referenced this pull request Jun 9, 2026
Review follow-up on #215 (comment #2). The miss-branch hint printed a bare
`export PATH=…` (fixes only the current shell) followed by `source ${rc}` on
an rc that did not yet contain the line — so neither command persisted the
fix, while the closing note implied ${rc} should hold it. The user is never
told to write the line into the rc. Rewrite the guidance per-shell:
- POSIX shells (zsh / bash / sh / dash): `echo '<export>' >> ${rc}` then
`source ${rc}` — one copy-pasteable step that fixes THIS terminal and every
new one.
- fish: `fish_add_path "…"` already persists (a universal var) AND applies to
the running shell, so drop the misleading `source ~/.config/fish/config.fish`.
Tests (install-cli.bats): the zsh miss-path now asserts the `echo … >> ~/.zshrc`
form; the fish case asserts no POSIX `export` and no `source`. bats 8/8 pass,
shellcheck --severity=error gate clean, bash -n clean.
NOTE: review comment #1 (fish fresh-shell probe using `command -v`, which fish's
`command` builtin lacks) is NOT addressed here — it needs verification on a real
fish and is tracked separately.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
saadqbal added a commit that referenced this pull request Jun 9, 2026
* feat(installer): self-verify CLI usability post-install with a shell-correct PATH fix (#738)
Step 5 installed the tracebloc CLI and then told the user "open a new
terminal so it's on your PATH" — without ever proving a fresh terminal
would actually find it. That is exactly the cli#61 failure mode (binary
lands in ~/.local/bin, which a brand-new shell doesn't have on PATH),
left undetected until the customer hits it. The installer is the last
place to catch it.
After the install attempt, self-verify and report precisely:
- Probe `command -v tracebloc` in BOTH a fresh login shell ("$SHELL" -lic)
and a non-login shell ("$SHELL" -ic) — they read different startup files
(~/.profile vs ~/.bashrc), and cli#61 was "works in my login shell,
missing in a plain `bash` subshell".
- If found: confirm via `tracebloc version` and print a VERIFIED verdict.
The canonical `tracebloc dataset push ./data` next step stays in the
summary's "What to do next" — not duplicated here.
- If a fresh shell would NOT find it: print the EXACT shell-correct fix
for the user's actual $SHELL (zsh→~/.zshrc, bash+linux→~/.bashrc,
bash+darwin→~/.bash_profile, fish→fish_add_path + ~/.config/fish/config.fish,
else ~/.profile), not a generic "open a new terminal".
Stays NON-FATAL by design: the client is already connected by Step 5, so
the verification always returns 0 and is hardened against the orchestrator's
`set -e`. Mirrored in install-k8s.ps1 (RefreshPath is the faithful
"fresh terminal" probe on Windows, since the CLI installer edits the
user-scope registry PATH).
Tests: extend install-cli.bats (verified-command success, actionable
shell-correct PATH hint on miss, fish-specific fix, non-fatal under
`set -e`) and mirror in install-k8s.Tests.ps1 (Test-TraceblocCli:
verified verdict, actionable hint, non-fatal when RefreshPath throws).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): make the #738 Windows CLI-verify Pester-safe on Linux CI
Two follow-ups so the Pester jobs go green (they were the only red checks
on #215; Pester is green on develop, so this PR introduced both):
- install-k8s.ps1: the new $TRACEBLOC_CLI_INSTALL_DIR ran Join-Path on
$env:LOCALAPPDATA at top level. The Pester suite dot-sources this script,
and on the Linux runner $env:LOCALAPPDATA is null — Join-Path throws on a
null -Path, aborting BeforeAll and failing the whole container (0/65).
Guard it; "" placeholder off Windows since the value is only used there.
- install-k8s.Tests.ps1: add a `function tracebloc { }` stub so
`Mock tracebloc` can bind. Pester v5 only mocks commands that already
exist (cf. the existing kubectl/docker/helm/k3d stubs); without it the
"fresh-shell success" test threw CommandNotFoundException — the lone
windows-latest failure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): make the #738 PATH-fix guidance actually persist
Review follow-up on #215 (comment #2). The miss-branch hint printed a bare
`export PATH=…` (fixes only the current shell) followed by `source ${rc}` on
an rc that did not yet contain the line — so neither command persisted the
fix, while the closing note implied ${rc} should hold it. The user is never
told to write the line into the rc. Rewrite the guidance per-shell:
- POSIX shells (zsh / bash / sh / dash): `echo '<export>' >> ${rc}` then
`source ${rc}` — one copy-pasteable step that fixes THIS terminal and every
new one.
- fish: `fish_add_path "…"` already persists (a universal var) AND applies to
the running shell, so drop the misleading `source ~/.config/fish/config.fish`.
Tests (install-cli.bats): the zsh miss-path now asserts the `echo … >> ~/.zshrc`
form; the fish case asserts no POSIX `export` and no `source`. bats 8/8 pass,
shellcheck --severity=error gate clean, bash -n clean.
NOTE: review comment #1 (fish fresh-shell probe using `command -v`, which fish's
`command` builtin lacks) is NOT addressed here — it needs verification on a real
fish and is tracked separately.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Asad Iqbal <asad.dsoft@gmail.com>
aptracebloc added a commit that referenced this pull request Jul 29, 2026
Bugbot + @saadqbal + a self code-review on client#458, all in slice-2's code:
- id -un everywhere (gate, _provision default, probe): $USER diverges from the
rootless daemon's user under su/cron, which wedged detection/provisioning (#1).
- Re-verify the uidmap helpers are usable (present AND setuid|cap_setuid) after
install, and return non-zero + warn (NOT error/exit) so run_prepare_host stays
best-effort while the installer sudo-path hard-fails via `|| error` (#2 + self-review).
- _idmap_helper_ok (common.sh): accept the setuid bit OR a cap_setuid filecap, so
Arch's `shadow`/pacman path isn't false-rejected (#3).
- Hand-off + run_prepare_host fallback compute a non-overlapping start via
_next_subid_start (honoring TB_SUBUID_FILE/TB_SUBGID_FILE), not hardcoded 100000
(#4 + self-review path-override).
- Hand-off command names the researcher (TB_PREPARE_USER=) — bare prepare-host
provisions nothing, so it would have looped back to the same hand-off (#5).
- Capture `usermod --help` before grepping — pipefail-safe (#6).
bats: id -un mocks, filecaps accept/reject, gate hand-off (names user + computed
start), _provision re-verify best-effort, run_prepare_host best-effort. R8 regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
aptracebloc added a commit that referenced this pull request Jul 29, 2026
Bugbot + @saadqbal + a self code-review on client#458, all in slice-2's code:
- id -un everywhere (gate, _provision default, probe): $USER diverges from the
rootless daemon's user under su/cron, which wedged detection/provisioning (#1).
- Re-verify the uidmap helpers are usable (present AND setuid|cap_setuid) after
install, and return non-zero + warn (NOT error/exit) so run_prepare_host stays
best-effort while the installer sudo-path hard-fails via `|| error` (#2 + self-review).
- _idmap_helper_ok (common.sh): accept the setuid bit OR a cap_setuid filecap, so
Arch's `shadow`/pacman path isn't false-rejected (#3).
- Hand-off + run_prepare_host fallback compute a non-overlapping start via
_next_subid_start (honoring TB_SUBUID_FILE/TB_SUBGID_FILE), not hardcoded 100000
(#4 + self-review path-override).
- Hand-off command names the researcher (TB_PREPARE_USER=) — bare prepare-host
provisions nothing, so it would have looped back to the same hand-off (#5).
- Capture `usermod --help` before grepping — pipefail-safe (#6).
bats: id -un mocks, filecaps accept/reject, gate hand-off (names user + computed
start), _provision re-verify best-effort, run_prepare_host best-effort. R8 regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
aptracebloc added a commit that referenced this pull request Jul 29, 2026
Bugbot + @saadqbal + a self code-review on client#458, all in slice-2's code:
- id -un everywhere (gate, _provision default, probe): $USER diverges from the
rootless daemon's user under su/cron, which wedged detection/provisioning (#1).
- Re-verify the uidmap helpers are usable (present AND setuid|cap_setuid) after
install, and return non-zero + warn (NOT error/exit) so run_prepare_host stays
best-effort while the installer sudo-path hard-fails via `|| error` (#2 + self-review).
- _idmap_helper_ok (common.sh): accept the setuid bit OR a cap_setuid filecap, so
Arch's `shadow`/pacman path isn't false-rejected (#3).
- Hand-off + run_prepare_host fallback compute a non-overlapping start via
_next_subid_start (honoring TB_SUBUID_FILE/TB_SUBGID_FILE), not hardcoded 100000
(#4 + self-review path-override).
- Hand-off command names the researcher (TB_PREPARE_USER=) — bare prepare-host
provisions nothing, so it would have looped back to the same hand-off (#5).
- Capture `usermod --help` before grepping — pipefail-safe (#6).
bats: id -un mocks, filecaps accept/reject, gate hand-off (names user + computed
start), _provision re-verify best-effort, run_prepare_host best-effort. R8 regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
aptracebloc added a commit that referenced this pull request Jul 29, 2026
…#1220) (#458)
* feat(install): Tier-1 subuid/subgid gate + prepare-host remediation
RFC 0001 #1220. Detect the one privileged residue a modern rootless host may
still need — a subordinate UID/GID range + the setuid uidmap helpers — and
either proceed (present), hand off to prepare-host (unprivileged), or perform
one announced touch (sudo available). Never blanket sudo, never an opaque
mid-install crash inside dockerd-rootless-setuptool.sh.
- probe.sh: _probe_subid_ranges (PROBE_SUBID) + _probe_uidmap_helpers
(PROBE_UIDMAP), set in run_host_probes (Linux only), plus audit rows on
the Tier-1 path.
- common.sh: shared pure parsers _subid_has_entry + _next_subid_start, used
by both the probe and the remediation (no duplication).
- setup-linux.sh: _ensure_subid_ranges gate (present / hand-off / one
announced sudo touch) called before install_rootless_docker;
_provision_subid_ranges (idempotent, non-overlapping block, usermod
--add-subuids with file-append fallback, uidmap install) shared by the
installer and run_prepare_host. Folds in slice-1's minimal uidmap check.
- Tests: probe.bats + setup-linux.bats. Manifest regenerated (R8).
Closestracebloc/backend#1220
Part of tracebloc/backend#1177 · Epic tracebloc/backend#1168
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address #458 review — subid gate/probe/hand-off robustness
Bugbot + @saadqbal + a self code-review on client#458, all in slice-2's code:
- id -un everywhere (gate, _provision default, probe): $USER diverges from the
rootless daemon's user under su/cron, which wedged detection/provisioning (#1).
- Re-verify the uidmap helpers are usable (present AND setuid|cap_setuid) after
install, and return non-zero + warn (NOT error/exit) so run_prepare_host stays
best-effort while the installer sudo-path hard-fails via `|| error` (#2 + self-review).
- _idmap_helper_ok (common.sh): accept the setuid bit OR a cap_setuid filecap, so
Arch's `shadow`/pacman path isn't false-rejected (#3).
- Hand-off + run_prepare_host fallback compute a non-overlapping start via
_next_subid_start (honoring TB_SUBUID_FILE/TB_SUBGID_FILE), not hardcoded 100000
(#4 + self-review path-override).
- Hand-off command names the researcher (TB_PREPARE_USER=) — bare prepare-host
provisions nothing, so it would have looped back to the same hand-off (#5).
- Capture `usermod --help` before grepping — pipefail-safe (#6).
bats: id -un mocks, filecaps accept/reject, gate hand-off (names user + computed
start), _provision re-verify best-effort, run_prepare_host best-effort. R8 regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(install): stub the subid gate in the Tier-1 rootless routing test
install_linux's Tier-1 branch now calls _ensure_subid_ranges (slice 2)
before install_rootless_docker; the routing test left it un-stubbed, so the
real gate hit the no-sudo hand-off and error()'d → install_linux returned
non-zero. Stub _ensure_subid_ranges (its own behavior is covered by the
dedicated gate tests) and assert it runs before daemon setup.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(install): scope set -o pipefail to a subshell (bats harness footgun)
Setting `set -o pipefail` in the @test body can leak into bats' own
post-test pipelines and fail the whole run with exit 1 even when every
test reports ok (no 'not ok'). Confine it to a subshell around the call
so the pipefail-safety assertion still holds without touching the harness.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): style guard — no bare curl in the prepare-host hint
The hand-off piped 'curl … | TB_PREPARE_USER=… bash', which breaks
check-style.sh's exemption for the canonical 'curl … | bash' one-liner
(the env var sits between the pipe and bash). Split into an 'export
TB_PREPARE_USER=…' line + the canonical piped one-liner — still names the
researcher, and passes the guard. Verified with scripts/check-style.sh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): fix the #458 red bats + 2 Bugbot bugs (newgidmap cap, write-failure)
Root cause of the "540 ok but exit 1" bats red: the probe.bats uidmap tests set
PATH="$bin" in the test body to hide system helpers, which also hides `rm` — so
bats-core 1.10+ can't run its own per-test cleanup ("rm: command not found") and
fails the whole run even though every test passes. Scope the hermetic PATH to a
subshell so it can't leak into bats' machinery. (Why develop was green + this was
so hard to see: these tests are new in slice 2, and the symptom is a clean pass
list with a non-zero exit.)
Two real Bugbot findings in the slice's own code:
- _idmap_helper_ok checked cap_setuid for BOTH helpers; newgidmap carries
cap_setgid (Arch filecaps) -> false-rejected. Map name->cap; fix the test mock
that masked it + add a wrong-cap regression test.
- _provision_subid_ranges printed success/returned 0 even when the usermod/tee
write failed (callers run it with set -e off) -> installer proceeds with no
range. Guard every write; warn + return 1 on failure. + a test.
Verified: probe.bats + setup-linux.bats EXIT 0 (0 not-ok, 0 rm-not-found) in a
faithful ubuntu 24.04 + bats 1.10 + non-root container. Rebased onto develop.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): 2 more Bugbot findings on #458 (apt hang, false zero-root message)
- _install_uidmap_pkg ran a bare `sudo apt-get install -y uidmap` under the
spinner — no needrestart/DEBIAN_FRONTEND env, no DPkg::Lock::Timeout, no
apt_wait_for_lock — so a headless Tier-1 install can hang on Ubuntu needrestart
or an apt-daily lock (#210 class). Reuse the repo's hardened PM_INSTALL
(populate via setup_pm, which Tier 1 skips) + apt_wait_for_lock.
- install_rootless_docker always printed "no administrator rights were used",
even after _ensure_subid_ranges performed an announced sudo touch on the
root/sudo_nopw path. The gate now sets TB_ROOTLESS_ADMIN_TOUCH and the summary
is honest on both the zero-root and one-admin-touch paths.
Tests: hardened-install assertion (NEEDRESTART_MODE + DPkg::Lock::Timeout) + a
success-message honesty test. Verified EXIT 0 (0 not-ok, 0 rm-errors) in the
faithful ubuntu 24.04 + bats 1.10 + non-root container.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): refresh the package index in _install_uidmap_pkg (Bugbot #458)
Completing the prior apt-hardening: _install_uidmap_pkg populated PM_INSTALL and
waited for the dpkg lock but never ran PM_UPDATE. On the Tier-1 path this is the
first package op, so an empty/stale index can't locate uidmap/shadow and the
install hard-stops. Run $PM_UPDATE (best-effort) first, matching the repo's other
install paths (setup-linux.sh:335/543). Test asserts the index refresh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
aptracebloc added a commit that referenced this pull request Jul 29, 2026
- _persist_docker_host: key idempotency off our own marker, not a bare
'DOCKER_HOST=' probe. The old probe also matched a user's own DOCKER_HOST
(remote/TCP), so we silently skipped persisting the rootless socket and new
shells kept hitting the wrong daemon. Now: our own line -> idempotent; a foreign
DOCKER_HOST -> left untouched + a warn to repoint it (Asad #2 + Bugbot #478, Medium).
- ensure_cluster_autostart: reset TB_DOCKER_AUTOSTART=0 in the rootless else-branch
(defensive; the is-enabled seed is already guarded off the rootless path) so the
honesty guarantee is local to the branch (Asad #1).
- install_rootless_docker: success line now reads "one or more one-time admin steps"
so it doesn't undercount when both the subuid and cgroup touches happen (Saqlain #1).
- Test: foreign DOCKER_HOST -> warns, no clobber, no double-write. manifest regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
aptracebloc added a commit that referenced this pull request Jul 29, 2026
… (#1221) (#478)
* feat(install): Tier-1 k3d-on-rootless + cgroup delegation + autostart (#1221)
Slice #1221 (RFC 0001 / #1177): make a rootless Tier-1 cluster actually usable,
all behind the opt-in TB_TIER1_ROOTLESS flag (default off until the spike's §5
host validation). With the flag unset every path below is a no-op and current
behavior is byte-for-byte unchanged.
- Shared _rootless_active predicate (common.sh) so cluster.sh + setup-linux.sh
can't drift on the flag pair.
- create_cluster targets the rootless socket (DOCKER_HOST); ensure_cluster_
autostart gets a user-scope branch (systemctl --user enable + loginctl
enable-linger, never `sudo systemctl enable docker`), and promises
reboot-survival only when BOTH succeed (honesty rule, #375/#458).
- cgroup v2 controller delegation drop-in (Delegate=cpu cpuset io memory pids):
privileged write + daemon-reload on root/sudo, or hand off to prepare-host
with the exact path+content when unprivileged. run_prepare_host writes it too
(system-wide -> covers the researcher).
- Carry-ins from #452/#458: scope-aware _configure_docker_proxy (user scope, no
sudo) so a proxy-only host's rootless daemon can pull rancher/k3s;
_set_tools_target installs user-space on rootless Tier 1 (no sudo-mv crash on
a true no-sudo host); persist DOCKER_HOST to the shell rc for new terminals.
14 new bats tests incl. flag-off regressions; shellcheck --severity=error clean;
manifest.sha256 regenerated (R8).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address Bugbot on #478 — rootless autostart seed + admin-touch msg
- ensure_cluster_autostart: don't seed TB_DOCKER_AUTOSTART from the SYSTEM
docker.service is-enabled check on the rootless path. The cluster runs on the
per-user rootless socket, so a system unit that happens to be enabled would seed
a false reboot promise the rootless branch then can't honestly retract. On
rootless the user-scope enable+linger are now the sole authority (Bugbot medium).
- install_rootless_docker: the TB_ROOTLESS_ADMIN_TOUCH success line no longer
hardcodes "subuid/subgid range" — _ensure_cgroup_delegation can set that flag
too, so it now names "host prerequisites (subuid/subgid range and/or cgroup
delegation)" (Bugbot low).
- Test: rootless + system docker.service enabled + user-enable fails => flag stays
0 (pins the seed-guard). manifest regenerated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address Asad + Saqlain review on #478
- _persist_docker_host: key idempotency off our own marker, not a bare
'DOCKER_HOST=' probe. The old probe also matched a user's own DOCKER_HOST
(remote/TCP), so we silently skipped persisting the rootless socket and new
shells kept hitting the wrong daemon. Now: our own line -> idempotent; a foreign
DOCKER_HOST -> left untouched + a warn to repoint it (Asad #2 + Bugbot #478, Medium).
- ensure_cluster_autostart: reset TB_DOCKER_AUTOSTART=0 in the rootless else-branch
(defensive; the is-enabled seed is already guarded off the rootless path) so the
honesty guarantee is local to the branch (Asad #1).
- install_rootless_docker: success line now reads "one or more one-time admin steps"
so it doesn't undercount when both the subuid and cgroup touches happen (Saqlain #1).
- Test: foreign DOCKER_HOST -> warns, no clobber, no double-write. manifest regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
LukasWodka added a commit that referenced this pull request Jul 30, 2026
* chore(chart): close values-schema gaps + drop dead override/code/docs (#963) (#457)
* chore(chart): close values-schema gaps + drop dead override/code/docs (#963)
Contract fixes for the client Helm chart (re-verified against develop at
chart v1.9.6; the #963 audit was taken at v1.8.4):
- values.schema.json: add the six live-but-unvalidated keys so bad values
fail `helm lint` instead of silently passing —
egressReachabilityCheck.enabled, ingestionAuthz.{allowed,serviceAccountName},
networkPolicy.training.enforcementProbeTimeoutSeconds, podTokenSigningSecret,
podTokenTtlSeconds. Types/defaults/constraints taken from values.yaml and
the templates that consume them. helm lint passes.
- ingestor subchart: remove the dead `image.repository` key — no template
ever rendered it (jobs-manager spawns from the parent chart's
images.ingestor.repository). Kept image.digest (live). README's air-gapped
override rows now point at the authoritative parent-chart path.
- README: drop the hardcoded chart version (said v1.3.5 while Chart.yaml is
1.9.6) and point to Chart.yaml / the releases page, so it can't drift again.
- Delete the unwired check_docker_arch_mac function + its bats test (no call
sites) and the orphaned docs/eks.md (referenced nowhere).
Part of tracebloc/backend#963.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(installer): regenerate manifest after common.sh trim + develop merge
The #963 chart-contract cleanup dropped 48 dead lines from
scripts/lib/common.sh, changing its sha; the installer manifest wasn't
regenerated, so the Static analysis gate (gen-manifest.sh --check) failed.
Merging develop also refreshed preflight.sh/install-k8s.ps1 hashes.
Regenerate scripts/manifest.sha256 to match the working tree.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Syed Saqlain <syedsaqlain@MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(release): package charts as the tag version + pre-releases skip gh-pages (#467)
* fix(release): package charts AS the tag version + pre-releases skip gh-pages
Incident 2026-07-29: the v1.9.7-rc.1 pre-release packaged the client chart
from Chart.yaml's plain 1.9.7 and pushed it into the public helm index as
a STABLE version -- customers running helm upgrade would have received
staging content (removed from the index by hand, tgz deleted). Two layers
now prevent it: (1) helm package --version/--app-version from the release
tag, so rc charts carry the -rc.N suffix helm's pre-release rules key on;
(2) pre-releases never run the gh-pages index steps at all -- FR consumes
the release assets (stamped installer / chart tgz), the index is a
customer surface reserved for finals.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: gh-pages gates key on verify's tag-derived prerelease, not the frozen event (Bugbot)
github.event.release.prerelease is an event-time snapshot: after verify
demotes a mis-marked release, it still reads false, so the demoted rc
would have entered the public index anyway. verify now outputs effective
prerelease-ness derived from the tag shape (the same strict rule the
demotion uses) and all three gh-pages steps gate on that output.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat: post-publish index-invariant job (manual leak catch -> CI)
After every release run: the public index must contain only stable-shaped
versions, and a prerelease run must not have indexed its own version.
Fails loudly; would have caught the 1.9.7 leak within a minute of it
happening instead of during manual FR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(install): rootless Docker core + Tier-1 routing (opt-in) (#1219) (#452)
* feat(install): rootless Docker core + Tier-1 routing (opt-in)
Add install_rootless_docker() and a Tier-1 early-branch in install_linux
so a modern-kernel host with no runtime and no root can install entirely
in user space (RFC 0001 Tier 1 — the RFC's primary path). Gated behind
opt-in TB_TIER1_ROOTLESS=1; with the flag unset a Tier-1 host falls
through to the legacy privileged flow unchanged (validated default).
- install_rootless_docker: uidmap-helper precondition (defers to
prepare-host #1178 when absent — never self-sudo), no-sudo install via
dockerd-rootless-setuptool.sh or get.docker.com/rootless, user-scoped
systemctl --user + loginctl enable-linger, DOCKER_HOST export with
XDG_RUNTIME_DIR fallback, single docker-info verify (no retry loop).
- Tier-1 branch mirrors the Tier-0 early-return. Tools still install via
sudo here (_set_tools_target keys no-sudo off Tier 0 only) — tightening
that for rootless Tier 1 is deferred to slice 3 (#1221).
- 6 bats cases; scripts/manifest.sha256 regenerated (R8).
Closestracebloc/backend#1219
Part of tracebloc/backend#1177 · Epic tracebloc/backend#1168
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: address Bugbot review on rootless Docker core (#452)
- Prepend ~/bin to PATH after the rootless install so this run's docker
info verify + later k3d/docker calls resolve the CLI the
get.docker.com/rootless fallback installs there (High).
- Bound the rootless `docker info` verify with a new shared _bounded
helper (timeout/gtimeout, mirrors probe.sh) so a wedged user daemon
can't hang a headless install (Medium).
- Guard the user-systemd bring-up under set -e: `systemctl --user … ||
true` (the bounded verify is the real gate) and `loginctl
enable-linger … || warn` (optional; fails on polkit-locked hosts even
when the daemon is up) (Medium).
Adds 2 bats cases (~/bin on PATH; systemd/linger failure falls through
to the verify). Manifest regenerated (R8).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: correct the uidmap remedy message (Bugbot #452)
The missing-uidmap error claimed prepare-host would install the uidmap
helpers, but run_prepare_host only sets up privileged Docker + the docker
group — it never installs uidmap. Point at the two honest remedies
instead: install the `uidmap` package directly (rootless then works), or
run prepare-host to set up Docker so the researcher installs at Tier 0
(no rootless needed). #1220 folds this into the shared subuid/subgid gate
and teaches prepare-host to install uidmap for real.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(install): TODO(#1221) — rootless daemon needs user-scoped proxy config
Bugbot on #452 flagged that install_rootless_docker never configures a
corporate proxy for the user-scoped dockerd (the #244 _configure_docker_proxy
is sudo/system-scoped and the Tier-1 early-return never reaches it), so k3d
pulls of rancher/k3s time out on proxy-only hosts. Deferred to #1221 (the
k3d-on-rootless-socket slice that owns the pulls); leaving a tracked TODO so
the follow-up adds the user-scoped drop-in.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: address Asad review nits on #452 — chmod no-op + misleading comment
- Drop the chmod +x on the rootless installer script: it runs via
`sh "$rootless_script"`, which ignores the exec bit.
- Reword the Tier-1 _install_userspace_tools comment: tools still
sudo-install on Tier 1 (only _persist_tools_on_path is no-sudo);
the comment previously implied otherwise.
The underlying _set_tools_target sudo-crash on no-sudo hosts and the
post-install DOCKER_HOST shell persistence are tracked to #1221.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: resolve user via id -un in install_rootless_docker (Saqlain review, #452)
$USER can be empty in headless / su / cron contexts (a Tier-1 target),
which would break `loginctl enable-linger` and the success line. Resolve
the user once via `id -un` (fallback $USER) and use it for the linger
call, its hint, and the success message. Matches the id-based robustness
DOCKER_HOST already uses. Happy-path bats now mocks `id -un` cleanly.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: add code-quality caller workflow (advisory) (#463)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Add pre-commit hooks (Layer 0, lint-only) (#465)
* Add pre-commit config (Layer 0, lint-only)
Lint-only on purpose: scripts/manifest.sha256 must keep matching the bytes
under scripts/, so no hook may rewrite files.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Document pre-commit setup in README
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs: add Bugbot resolve-and-reply team norm to .cursor/BUGBOT.md (#464)
Part of tracebloc/backend#1308
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* ci: cover scripts/resolve-ingestor-digest.sh in CI shellcheck (#466)
* ci: lint scripts/resolve-ingestor-digest.sh in CI shellcheck (was never linted)
Both CI shellcheck invocations enumerate files explicitly and both omitted
this script. Verified clean against shellcheck --severity=error --shell=bash
0.11.0 before adding. The pre-commit hook from #465 already covers it
locally; this closes the same gap on the CI side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* ci: lint scripts/resolve-ingestor-digest.sh in CI shellcheck (was never linted)
Both CI shellcheck invocations enumerate files explicitly and both omitted
this script. Verified clean against shellcheck --severity=error --shell=bash
0.11.0 before adding. The pre-commit hook from #465 already covers it
locally; this closes the same gap on the CI side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): honest cosign bootstrap download + translated DISM failures (#468) (#469)
The v1.9.7-rc.1 FR killed a healthy install: PS 5.1's progress overlay
throttled the 17 MB pinned-cosign fetch to ~4.5 min of dead silence and
the window read as frozen.
- silence the PS 5.1 progress overlay in Get-WithRetry/Get-Optional
(function-local, auto-reverts) - the classic 10-50x IWR speedup
- run the cosign fetch in a background job with a dim liveness tick
(Wait-JobWithTicks / Get-OptionalWithTicks; cwd pinned per #409,
TLS 1.2 re-applied in the fresh process), expectation lines before,
elapsed + checksum-verified confirmation after
- ASCII-only string literals in both installers: the release asset is
served without a charset so PS 5.1's irm decodes UTF-8 source as
Latin-1 before iex, and BOM-less -File reads are ANSI - literal
em-dashes/ellipses reached customers as mojibake. Locked in by a
tokenizer-based Pester test (which also caught the -Help here-string).
- Enable-OneVirtFeature: translate DISM's raw COMException (feature
package absent on Server SKUs vs enable failure) and stop demanding
a reboot for a feature that never enabled (old code sent Server
users into a reboot->re-run->same-error loop)
Pester: 212 passed / 0 failed locally (pwsh 7.5, Pester 5.7.1).
PSScriptAnalyzer: 0 errors. manifest.sha256 regenerated.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(installer): trust the corporate MITM CA in the k3d nodes (#424) (#453)
* fix(installer): trust the corporate MITM CA in the k3d nodes (#424)
Proxy REACHABILITY reaches the nodes, but on a TLS-inspecting (break-and-
inspect) network the nodes still don't TRUST the corporate CA, so every in-node
containerd pull (rancher/k3s, ghcr.io, tracebloc images) fails x509 — then
masked (helm runs without --wait) into a root-cause-free "an image couldn't be
pulled." Enterprise/hospital archetype, all three OSes.
- Inject the CA at create time: when TRACEBLOC_CA_BUNDLE (or CURL_CA_BUNDLE) is
set, mount the bundle into every k3d node and write a registries.yaml pointing
containerd at it per-registry (docker.io, registry-1.docker.io, ghcr.io), via
the same --config/create path that already carries proxy env. Parity across
scripts/lib/cluster.sh (Linux/macOS) and install-k8s.ps1 (Windows). A CA var
set but unreadable fails loudly instead of silently skipping.
- Name the env var where the user hits the wall: the TLS-interception preflight
hint (both OSes), docs/INSTALL.md, and the PS -Help env-var list.
- CA-aware diagnosis: detect x509 / "certificate signed by unknown authority"
pull events and report a dedicated image_pull_ca state — "the cluster does not
trust your network's TLS-inspection CA" + the exact remedy — instead of the
generic pull error. Mirrored in summary.sh and Print-Summary.
- New check-drift.sh parity check (_drift_ca_trust) so neither installer can drop
the CA wiring for the other's OS.
Tests: +8 cluster.bats, +3 summary.bats, +2 check-drift.bats, +8 Pester.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): CA-trust hardening — no fail-open, bounded events, verify CA readable (Bugbot #424)
Three Bugbot findings on #424:
- _write_k3d_registries_config failed open: on mktemp failure it returned success
with no path, so create still mounted the CA and logged "nodes trust it" but
dropped --registry-config → containerd never got ca_file, x509 pulls still fail
while the operator thinks it's fixed. Now returns non-zero; the caller hard-errors
(CA was supplied, so we refuse to proceed without wiring it in).
- PS Get-NotReadyState `kubectl get events` had no --request-timeout (the bash path
does) — on a wedged/proxy-misrouted API, classification could hang. Added
--request-timeout=5s to match _diagnose_not_ready.
- PS Resolve-CaBundle only checked existence (Test-Path), not readability, so an
unreadable CA passed on Windows but bash (-r) hard-fails. Added an OpenRead probe
so both fail the same way, up front.
Tests: cluster.bats +mktemp-failure + unwritable-registries-hard-error;
install-k8s.Tests.ps1 +unreadable-CA (Unix) + events --request-timeout assertion.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): errexit-safe CA-resolve capture + drift check ignores comments (Bugbot #424 r2)
Two round-2 Bugbot findings:
- Under `set -euo pipefail`, `ca_bundle="$(_resolve_ca_bundle)"; ca_rc=$?` exited on
the rc-2 (unreadable/missing CA) BEFORE ca_rc/error ran — operators got a bare
exit instead of the "can't be read" guidance. Capture with `|| ca_rc=$?` so
errexit doesn't fire and the guidance prints.
- _drift_ca_trust whole-file grep matched tokens in comments (e.g. --registry-config
appears in a comment above the real line), so deleting the functional wiring could
still pass. Strip comment lines first (matches the execute-gate / preflight-host
checks), no grep -q under pipefail.
Tests: cluster.bats +errexit-safe-capture; check-drift.bats +comment-only-token drift.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): TLS-preflight hint names the right var per layer/OS (Bugbot #424 r3)
The hint claimed TRACEBLOC_CA_BUNDLE makes "the host AND the k3d nodes" trust the
CA, but the host connectivity checks use curl_secure / Invoke-WebRequest, which
read CURL_CA_BUNDLE / the system trust store — not TRACEBLOC_CA_BUNDLE (that var
only reaches the nodes via _resolve_ca_bundle). Following the hint literally left
host preflight TLS failures unchanged. Corrected, no behaviour change:
- bash: CURL_CA_BUNDLE fixes these host checks AND the nodes; TRACEBLOC_CA_BUNDLE
is nodes-only; or add the CA to the system trust store.
- Windows: import the CA into the cert store for the host checks (Invoke-WebRequest
uses the store, not an env var); TRACEBLOC_CA_BUNDLE/CURL_CA_BUNDLE cover the nodes.
(Reworded to avoid a bare lowercase `curl` that the curl_secure style guard flags.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): apply CA on cluster REUSE path — warn + recreate guidance (Bugbot #424 r4)
The image_pull_ca remedy said "set the CA and re-run", but CA trust is baked in
only at fresh create; a re-run reuses the existing cluster and never mounts the CA
or passes --registry-config, so the x509 pulls persisted. Mirror the existing
proxy handling (baked-at-create → warn on reuse):
- bash _check_existing_cluster_ca (called from _handle_existing_cluster): warns when
a CA bundle is set but the reused server container lacks the CA mount.
- ps1 New-K3dCluster reuse block: same check via docker inspect mounts.
- both image_pull_ca remedies now say to `k3d cluster delete <name>` first, then
re-run with the CA (CA, like proxy, can't be added to a running cluster).
Tests: cluster.bats +3 (no-CA no-op / CA-but-missing-mount warns / mount-present silent).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): add auth.docker.io to the CA registries config (Bugbot #424 r5)
The registries.yaml ca_file entries covered docker.io / registry-1.docker.io /
ghcr.io, but Docker Hub pulls also TLS-handshake with auth.docker.io for bearer
tokens — so on a break-and-inspect network containerd still rejected the
intercepted cert there even with the CA mounted. #416 already probes auth.docker.io
at preflight; the CA registries list now matches. Added to TB_CA_REGISTRIES and
$TbCaRegistries; registries.yaml test counts 3 -> 4.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#424): _resolve_ca_bundle rejects a directory, not just unreadable paths
A directory of PEMs is readable (-r) but would bind-mount over the single
node ca_file path and containerd can't read it — the silent 'looks applied
but still x509' case. Require a regular file (-f), mirroring the PS
Resolve-CaBundle -PathType Leaf check. Adds a directory-reject bats case.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#424): exact whole-line match for CA mount detection (Bugbot)
_check_existing_cluster_ca used a substring test on docker mount
destinations, so a longer path embedding /etc/ssl/certs/tracebloc-mitm-ca.crt
(e.g. …crt.bak) would be treated as the CA mount and skip the recreate
warning while containerd still x509-fails. Switch to grep -qxF (exact
whole-line), matching the PS anchored regex. Adds a substring-embed test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#424): scope x509 classification to the pull-failure event (Asad)
_diagnose_not_ready / Get-NotReadyState flagged image_pull_ca on ANY x509
event in the namespace, so a stale/unrelated x509 event (e.g. a FailedMount)
could misdirect the user into a needless delete+recreate. Filter events to
the image-pull failure lines (failed to pull / ErrImagePull) before testing
x509, in both bash and PS. Adds an unrelated-x509 test to each side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(chart): perIngestionTables — RFC-0003 D16 enablement knob (backend#1205) (#472)
* feat(chart): perIngestionTables — the RFC-0003 D16 enablement knob (backend#1204/#1205)
values.perIngestionTables (default false, schema-typed) renders
PER_INGESTION_TABLES=1 onto the jobs-manager, which forwards it into
every ingestion Job it spawns (client-runtime companion PR). Flip per
environment, dev first, only once that environment's backend + engine
images + jobs-manager carry the merged D-series. Default installs
render byte-identically (conditional block; unit tests pin both sides).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(values): own banner for perIngestionTables — it is not part of the authz section (review)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* chore: clear house-rules findings (#470)
Fix every finding the shared org checker (tracebloc/.github
scripts/house-rules.sh) reports at develop HEAD: missing curl
timeouts/TLS floors, plus (cli) a missing pipefail. Waivers only where
the finding is a documented false positive. Part of tracebloc/backend#1303.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(install): Tier-1 subuid/subgid gate + prepare-host remediation (#1220) (#458)
* feat(install): Tier-1 subuid/subgid gate + prepare-host remediation
RFC 0001 #1220. Detect the one privileged residue a modern rootless host may
still need — a subordinate UID/GID range + the setuid uidmap helpers — and
either proceed (present), hand off to prepare-host (unprivileged), or perform
one announced touch (sudo available). Never blanket sudo, never an opaque
mid-install crash inside dockerd-rootless-setuptool.sh.
- probe.sh: _probe_subid_ranges (PROBE_SUBID) + _probe_uidmap_helpers
(PROBE_UIDMAP), set in run_host_probes (Linux only), plus audit rows on
the Tier-1 path.
- common.sh: shared pure parsers _subid_has_entry + _next_subid_start, used
by both the probe and the remediation (no duplication).
- setup-linux.sh: _ensure_subid_ranges gate (present / hand-off / one
announced sudo touch) called before install_rootless_docker;
_provision_subid_ranges (idempotent, non-overlapping block, usermod
--add-subuids with file-append fallback, uidmap install) shared by the
installer and run_prepare_host. Folds in slice-1's minimal uidmap check.
- Tests: probe.bats + setup-linux.bats. Manifest regenerated (R8).
Closestracebloc/backend#1220
Part of tracebloc/backend#1177 · Epic tracebloc/backend#1168
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address #458 review — subid gate/probe/hand-off robustness
Bugbot + @saadqbal + a self code-review on client#458, all in slice-2's code:
- id -un everywhere (gate, _provision default, probe): $USER diverges from the
rootless daemon's user under su/cron, which wedged detection/provisioning (#1).
- Re-verify the uidmap helpers are usable (present AND setuid|cap_setuid) after
install, and return non-zero + warn (NOT error/exit) so run_prepare_host stays
best-effort while the installer sudo-path hard-fails via `|| error` (#2 + self-review).
- _idmap_helper_ok (common.sh): accept the setuid bit OR a cap_setuid filecap, so
Arch's `shadow`/pacman path isn't false-rejected (#3).
- Hand-off + run_prepare_host fallback compute a non-overlapping start via
_next_subid_start (honoring TB_SUBUID_FILE/TB_SUBGID_FILE), not hardcoded 100000
(#4 + self-review path-override).
- Hand-off command names the researcher (TB_PREPARE_USER=) — bare prepare-host
provisions nothing, so it would have looped back to the same hand-off (#5).
- Capture `usermod --help` before grepping — pipefail-safe (#6).
bats: id -un mocks, filecaps accept/reject, gate hand-off (names user + computed
start), _provision re-verify best-effort, run_prepare_host best-effort. R8 regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(install): stub the subid gate in the Tier-1 rootless routing test
install_linux's Tier-1 branch now calls _ensure_subid_ranges (slice 2)
before install_rootless_docker; the routing test left it un-stubbed, so the
real gate hit the no-sudo hand-off and error()'d → install_linux returned
non-zero. Stub _ensure_subid_ranges (its own behavior is covered by the
dedicated gate tests) and assert it runs before daemon setup.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(install): scope set -o pipefail to a subshell (bats harness footgun)
Setting `set -o pipefail` in the @test body can leak into bats' own
post-test pipelines and fail the whole run with exit 1 even when every
test reports ok (no 'not ok'). Confine it to a subshell around the call
so the pipefail-safety assertion still holds without touching the harness.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): style guard — no bare curl in the prepare-host hint
The hand-off piped 'curl … | TB_PREPARE_USER=… bash', which breaks
check-style.sh's exemption for the canonical 'curl … | bash' one-liner
(the env var sits between the pipe and bash). Split into an 'export
TB_PREPARE_USER=…' line + the canonical piped one-liner — still names the
researcher, and passes the guard. Verified with scripts/check-style.sh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): fix the #458 red bats + 2 Bugbot bugs (newgidmap cap, write-failure)
Root cause of the "540 ok but exit 1" bats red: the probe.bats uidmap tests set
PATH="$bin" in the test body to hide system helpers, which also hides `rm` — so
bats-core 1.10+ can't run its own per-test cleanup ("rm: command not found") and
fails the whole run even though every test passes. Scope the hermetic PATH to a
subshell so it can't leak into bats' machinery. (Why develop was green + this was
so hard to see: these tests are new in slice 2, and the symptom is a clean pass
list with a non-zero exit.)
Two real Bugbot findings in the slice's own code:
- _idmap_helper_ok checked cap_setuid for BOTH helpers; newgidmap carries
cap_setgid (Arch filecaps) -> false-rejected. Map name->cap; fix the test mock
that masked it + add a wrong-cap regression test.
- _provision_subid_ranges printed success/returned 0 even when the usermod/tee
write failed (callers run it with set -e off) -> installer proceeds with no
range. Guard every write; warn + return 1 on failure. + a test.
Verified: probe.bats + setup-linux.bats EXIT 0 (0 not-ok, 0 rm-not-found) in a
faithful ubuntu 24.04 + bats 1.10 + non-root container. Rebased onto develop.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): 2 more Bugbot findings on #458 (apt hang, false zero-root message)
- _install_uidmap_pkg ran a bare `sudo apt-get install -y uidmap` under the
spinner — no needrestart/DEBIAN_FRONTEND env, no DPkg::Lock::Timeout, no
apt_wait_for_lock — so a headless Tier-1 install can hang on Ubuntu needrestart
or an apt-daily lock (#210 class). Reuse the repo's hardened PM_INSTALL
(populate via setup_pm, which Tier 1 skips) + apt_wait_for_lock.
- install_rootless_docker always printed "no administrator rights were used",
even after _ensure_subid_ranges performed an announced sudo touch on the
root/sudo_nopw path. The gate now sets TB_ROOTLESS_ADMIN_TOUCH and the summary
is honest on both the zero-root and one-admin-touch paths.
Tests: hardened-install assertion (NEEDRESTART_MODE + DPkg::Lock::Timeout) + a
success-message honesty test. Verified EXIT 0 (0 not-ok, 0 rm-errors) in the
faithful ubuntu 24.04 + bats 1.10 + non-root container.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): refresh the package index in _install_uidmap_pkg (Bugbot #458)
Completing the prior apt-hardening: _install_uidmap_pkg populated PM_INSTALL and
waited for the dpkg lock but never ran PM_UPDATE. On the Tier-1 path this is the
first package op, so an empty/stale index can't locate uidmap/shadow and the
install hard-stops. Run $PM_UPDATE (best-effort) first, matching the repo's other
install paths (setup-linux.sh:335/543). Test asserts the index refresh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(installer): silence the PS 5.1 progress throttle in install-k8s.ps1 downloads (#468 follow-up) (#471)
Same class as the bootstrap fix in #469: PS 5.1's progress overlay
throttles Invoke-WebRequest 10-50x and reads like a hang. One
function-local $ProgressPreference in Invoke-WithRetry covers every
fetch scriptblock it drives (dynamic scoping) - winget msixbundle,
Docker Desktop fallback, kubectl, k3d, helm, GPU plugin yaml, and the
version resolvers.
Honest-progress expectation lines (sizes measured today via HEAD):
Docker Desktop ~600 MB, winget ~200 MB, kubectl ~60 MB, k3d ~25 MB,
helm ~20 MB - all cold-path only, silent on warm re-runs.
Pester: 205 passed / 0 failed locally. PSSA: 0 errors.
manifest.sha256 regenerated.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: shujaat hasan <shujaat@tracebloc.io>
* fix(installer): trust the corporate CA in the Docker daemon for k3d runtime image pulls (#474) (#475)
* fix(#474): detect host Docker daemon x509 at cluster-create + document daemon CA trust
k3d pulls its own runtime images (rancher/k3s, k3d-tools, k3d-proxy) with the
HOST Docker daemon, which doesn't use the in-node CA trust from #424. On a
TLS-inspecting network that pull can x509-fail during 'k3d cluster create',
before any node boots — so the post-create diagnosis never classifies it.
- bash: _host_ca_create_hint() detects x509 in the k3d create output and prints
a platform-aware remedy (Linux system trust store vs Docker Desktop VM);
wired into _create_new_cluster's failure path.
- PS: Write-HostCaCreateHint() mirrors it (Windows Trusted Root store), wired
before the generic create failure.
- docs/INSTALL.md: document trusting the CA in the daemon itself (Linux /
Docker Desktop).
- check-drift.sh: enforce both installers keep the host-CA hint (parity).
- Tests: bats (Linux/macOS branches + silent-on-no-x509) + Pester + drift.
Closes#474
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): distro-aware Linux CA remedy + actionable Docker Desktop for Linux (Bugbot)
- Linux native-Docker remedy now covers both Debian/Ubuntu (update-ca-certificates)
and RHEL/Fedora (update-ca-trust), not just the Debian path — the installer
supports RHEL hosts where the Debian commands fail.
- Docker Desktop for Linux now has an actionable step (trust in the system store,
restart Docker Desktop) instead of a dangling reference to a step only printed
on the macOS branch.
- docs/INSTALL.md updated to match. bats Linux test asserts both distro paths.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): cover Colima runtime in the macOS host-CA remedy (Bugbot)
Headless macOS installs use Colima (_install_docker_colima), a Lima VM that
does not read the macOS keychain — so the 'trust it in the keychain + restart
Docker Desktop' remedy was wrong for those hosts. The macOS branch now also
gives the Colima path (add the CA inside the VM via 'colima ssh', then
'colima restart'). docs/INSTALL.md + macOS bats test updated to match.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#474): isolate drift negative fixtures so a missing new token can't mask them (Bugbot)
Adding _host_ca_create_hint / Write-HostCaCreateHint as required _drift_ca_trust
tokens meant the older negative fixtures (missing registry-config, comment-only
registry-config) could pass just because the new token was also absent — so the
comment-strip case no longer uniquely proved comment-stripping still works. Each
negative fixture now carries ALL other required tokens and omits/comments only
the one under test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): surface host-CA hint on the PS create-timeout path too (Bugbot parity)
The PowerShell create-timeout branch exited via Err without calling
Write-HostCaCreateHint (and deleted the k3d logs first), so a TLS-inspected
host pull that logs x509 then hangs to the deadline gave Windows operators a
raw timeout with no certlm.msc CA guidance — while bash runs _host_ca_create_hint
on its timeout fall-through. Capture the full create output before deleting the
logs and call the hint before the timeout Err. Adds a parity regression test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): use a herestring in _host_ca_create_hint to survive pipefail (Asad)
printf '%s' "$out" | grep -qiE ... could swallow the hint under set -o pipefail:
grep -q closes the pipe on first match, so for output past the ~64KB pipe buffer
(reachable on the timeout path, which passes the full logs) printf takes SIGPIPE,
the pipeline exits non-zero, and `|| return 0` bails even though x509 matched.
Feed grep via a herestring (no pipe, no SIGPIPE). Adds a >64KB-under-pipefail
regression test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(install): Tier-1 k3d-on-rootless + cgroup delegation + autostart (#1221) (#478)
* feat(install): Tier-1 k3d-on-rootless + cgroup delegation + autostart (#1221)
Slice #1221 (RFC 0001 / #1177): make a rootless Tier-1 cluster actually usable,
all behind the opt-in TB_TIER1_ROOTLESS flag (default off until the spike's §5
host validation). With the flag unset every path below is a no-op and current
behavior is byte-for-byte unchanged.
- Shared _rootless_active predicate (common.sh) so cluster.sh + setup-linux.sh
can't drift on the flag pair.
- create_cluster targets the rootless socket (DOCKER_HOST); ensure_cluster_
autostart gets a user-scope branch (systemctl --user enable + loginctl
enable-linger, never `sudo systemctl enable docker`), and promises
reboot-survival only when BOTH succeed (honesty rule, #375/#458).
- cgroup v2 controller delegation drop-in (Delegate=cpu cpuset io memory pids):
privileged write + daemon-reload on root/sudo, or hand off to prepare-host
with the exact path+content when unprivileged. run_prepare_host writes it too
(system-wide -> covers the researcher).
- Carry-ins from #452/#458: scope-aware _configure_docker_proxy (user scope, no
sudo) so a proxy-only host's rootless daemon can pull rancher/k3s;
_set_tools_target installs user-space on rootless Tier 1 (no sudo-mv crash on
a true no-sudo host); persist DOCKER_HOST to the shell rc for new terminals.
14 new bats tests incl. flag-off regressions; shellcheck --severity=error clean;
manifest.sha256 regenerated (R8).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address Bugbot on #478 — rootless autostart seed + admin-touch msg
- ensure_cluster_autostart: don't seed TB_DOCKER_AUTOSTART from the SYSTEM
docker.service is-enabled check on the rootless path. The cluster runs on the
per-user rootless socket, so a system unit that happens to be enabled would seed
a false reboot promise the rootless branch then can't honestly retract. On
rootless the user-scope enable+linger are now the sole authority (Bugbot medium).
- install_rootless_docker: the TB_ROOTLESS_ADMIN_TOUCH success line no longer
hardcodes "subuid/subgid range" — _ensure_cgroup_delegation can set that flag
too, so it now names "host prerequisites (subuid/subgid range and/or cgroup
delegation)" (Bugbot low).
- Test: rootless + system docker.service enabled + user-enable fails => flag stays
0 (pins the seed-guard). manifest regenerated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address Asad + Saqlain review on #478
- _persist_docker_host: key idempotency off our own marker, not a bare
'DOCKER_HOST=' probe. The old probe also matched a user's own DOCKER_HOST
(remote/TCP), so we silently skipped persisting the rootless socket and new
shells kept hitting the wrong daemon. Now: our own line -> idempotent; a foreign
DOCKER_HOST -> left untouched + a warn to repoint it (Asad #2 + Bugbot #478, Medium).
- ensure_cluster_autostart: reset TB_DOCKER_AUTOSTART=0 in the rootless else-branch
(defensive; the is-enabled seed is already guarded off the rootless path) so the
honesty guarantee is local to the branch (Asad #1).
- install_rootless_docker: success line now reads "one or more one-time admin steps"
so it doesn't undercount when both the subuid and cgroup touches happen (Saqlain #1).
- Test: foreign DOCKER_HOST -> warns, no clobber, no double-write. manifest regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(installer): failures surface the real error + log path + support-bundle hint (#423) (#476)
* fix(#423): failures surface the real error + log path + support-bundle hint
Fatal errors printed a generic red line while the actionable detail (k3d/helm
stderr) went only to the transcript, and the log path itself was never shown on
screen. Now:
- Err gains an optional $Detail param; Get-ErrDetailLines (pure, unit-tested)
renders the last ~5 non-empty output lines + the log path + a '-Diagnose'
next-step hint, appended to EVERY fatal error.
- Cluster-create failure passes k3d's stdout/stderr so the real reason (image
pull / proxy / port / WSL) shows on screen — the motivating case.
- Helm repo-add / reconcile / install failures pass helm's output via $Detail
instead of embedding it (no more duplicated log-path text).
- Install log path is announced up front in the banner (was log-only before).
Closes#423
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): force array enumeration of Err detail lines (Bugbot, defensive)
Bugbot flagged that a single-line Get-ErrDetailLines return (no detail + no
LOG_FILE, e.g. a Confirm-Config failure before Start-InstallLog) unwraps to a
scalar string. The foreach statement already iterates a scalar once (verified:
it prints the whole line, not per-character), so the reported char-splitting
does not reproduce -- but wrap the enumeration in @(...) to make that
unambiguous and future-proof. Adds a regression test asserting the single-line
case stays one intact line.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): strip PS 5.1 ErrorRecord chrome from the failure excerpt (Bugbot)
helm failures arrive as `native 2>&1 | Out-String`; on Windows PowerShell 5.1
that wraps stderr in ErrorRecord chrome (the `At <file>:<n> char:<n>` position
line plus the `+ ...` / `+ CategoryInfo` / `+ FullyQualifiedErrorId` block).
Get-ErrDetailLines kept only the last 5 non-empty lines, so the excerpt was all
chrome and the real `Error:` line dropped out -- a regression from the previous
full-message dump. Filter those chrome lines before taking the window so the
actual error survives. Adds a regression test simulating the 5.1 rendering.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): drop duplicate inline log-path hints (Bugbot)
Err now always prints the log path via Get-ErrDetailLines, so the k3d
spawn-failure and create-timeout paths that still Hint "Full log:" right before
Err printed it twice. Remove those inline hints; Err is the single source. Adds
a guard test asserting no inline 'Full log:' hints remain in the installer.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): put stderr last in the create-failure Err detail (Asad)
Get-ErrDetailLines keeps the LAST 5 non-empty lines, so with detail ordered
stderr-then-stdout any k3d stdout tail could crowd the real stderr reason
(FATA/x509/port) out of the excerpt. Order it stdout-then-stderr so the stderr
tail survives the window; also matches the Write-HostCaCreateHint order just
above.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Syed Is Saqlain <saqlain.syed007@gmail.com>
Co-authored-by: Syed Saqlain <syedsaqlain@MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Arturo Peroni <arturo@tracebloc.io>
Co-authored-by: shujaat_tracebloc <153823837+shujaatTracebloc@users.noreply.github.com>
Co-authored-by: shujaat hasan <shujaat@tracebloc.io>
Co-authored-by: tracebloc-release-train[bot] <309815517+tracebloc-release-train[bot]@users.noreply.github.com>
LukasWodka added a commit that referenced this pull request Aug 15, 2026
…pying it (#723)
* release-train: staging -> main (#495)
* chore(chart): close values-schema gaps + drop dead override/code/docs (#963) (#457)
* chore(chart): close values-schema gaps + drop dead override/code/docs (#963)
Contract fixes for the client Helm chart (re-verified against develop at
chart v1.9.6; the #963 audit was taken at v1.8.4):
- values.schema.json: add the six live-but-unvalidated keys so bad values
fail `helm lint` instead of silently passing —
egressReachabilityCheck.enabled, ingestionAuthz.{allowed,serviceAccountName},
networkPolicy.training.enforcementProbeTimeoutSeconds, podTokenSigningSecret,
podTokenTtlSeconds. Types/defaults/constraints taken from values.yaml and
the templates that consume them. helm lint passes.
- ingestor subchart: remove the dead `image.repository` key — no template
ever rendered it (jobs-manager spawns from the parent chart's
images.ingestor.repository). Kept image.digest (live). README's air-gapped
override rows now point at the authoritative parent-chart path.
- README: drop the hardcoded chart version (said v1.3.5 while Chart.yaml is
1.9.6) and point to Chart.yaml / the releases page, so it can't drift again.
- Delete the unwired check_docker_arch_mac function + its bats test (no call
sites) and the orphaned docs/eks.md (referenced nowhere).
Part of tracebloc/backend#963.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(installer): regenerate manifest after common.sh trim + develop merge
The #963 chart-contract cleanup dropped 48 dead lines from
scripts/lib/common.sh, changing its sha; the installer manifest wasn't
regenerated, so the Static analysis gate (gen-manifest.sh --check) failed.
Merging develop also refreshed preflight.sh/install-k8s.ps1 hashes.
Regenerate scripts/manifest.sha256 to match the working tree.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Syed Saqlain <syedsaqlain@MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(release): package charts as the tag version + pre-releases skip gh-pages (#467)
* fix(release): package charts AS the tag version + pre-releases skip gh-pages
Incident 2026-07-29: the v1.9.7-rc.1 pre-release packaged the client chart
from Chart.yaml's plain 1.9.7 and pushed it into the public helm index as
a STABLE version -- customers running helm upgrade would have received
staging content (removed from the index by hand, tgz deleted). Two layers
now prevent it: (1) helm package --version/--app-version from the release
tag, so rc charts carry the -rc.N suffix helm's pre-release rules key on;
(2) pre-releases never run the gh-pages index steps at all -- FR consumes
the release assets (stamped installer / chart tgz), the index is a
customer surface reserved for finals.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: gh-pages gates key on verify's tag-derived prerelease, not the frozen event (Bugbot)
github.event.release.prerelease is an event-time snapshot: after verify
demotes a mis-marked release, it still reads false, so the demoted rc
would have entered the public index anyway. verify now outputs effective
prerelease-ness derived from the tag shape (the same strict rule the
demotion uses) and all three gh-pages steps gate on that output.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat: post-publish index-invariant job (manual leak catch -> CI)
After every release run: the public index must contain only stable-shaped
versions, and a prerelease run must not have indexed its own version.
Fails loudly; would have caught the 1.9.7 leak within a minute of it
happening instead of during manual FR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(install): rootless Docker core + Tier-1 routing (opt-in) (#1219) (#452)
* feat(install): rootless Docker core + Tier-1 routing (opt-in)
Add install_rootless_docker() and a Tier-1 early-branch in install_linux
so a modern-kernel host with no runtime and no root can install entirely
in user space (RFC 0001 Tier 1 — the RFC's primary path). Gated behind
opt-in TB_TIER1_ROOTLESS=1; with the flag unset a Tier-1 host falls
through to the legacy privileged flow unchanged (validated default).
- install_rootless_docker: uidmap-helper precondition (defers to
prepare-host #1178 when absent — never self-sudo), no-sudo install via
dockerd-rootless-setuptool.sh or get.docker.com/rootless, user-scoped
systemctl --user + loginctl enable-linger, DOCKER_HOST export with
XDG_RUNTIME_DIR fallback, single docker-info verify (no retry loop).
- Tier-1 branch mirrors the Tier-0 early-return. Tools still install via
sudo here (_set_tools_target keys no-sudo off Tier 0 only) — tightening
that for rootless Tier 1 is deferred to slice 3 (#1221).
- 6 bats cases; scripts/manifest.sha256 regenerated (R8).
Closestracebloc/backend#1219
Part of tracebloc/backend#1177 · Epic tracebloc/backend#1168
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: address Bugbot review on rootless Docker core (#452)
- Prepend ~/bin to PATH after the rootless install so this run's docker
info verify + later k3d/docker calls resolve the CLI the
get.docker.com/rootless fallback installs there (High).
- Bound the rootless `docker info` verify with a new shared _bounded
helper (timeout/gtimeout, mirrors probe.sh) so a wedged user daemon
can't hang a headless install (Medium).
- Guard the user-systemd bring-up under set -e: `systemctl --user … ||
true` (the bounded verify is the real gate) and `loginctl
enable-linger … || warn` (optional; fails on polkit-locked hosts even
when the daemon is up) (Medium).
Adds 2 bats cases (~/bin on PATH; systemd/linger failure falls through
to the verify). Manifest regenerated (R8).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: correct the uidmap remedy message (Bugbot #452)
The missing-uidmap error claimed prepare-host would install the uidmap
helpers, but run_prepare_host only sets up privileged Docker + the docker
group — it never installs uidmap. Point at the two honest remedies
instead: install the `uidmap` package directly (rootless then works), or
run prepare-host to set up Docker so the researcher installs at Tier 0
(no rootless needed). #1220 folds this into the shared subuid/subgid gate
and teaches prepare-host to install uidmap for real.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(install): TODO(#1221) — rootless daemon needs user-scoped proxy config
Bugbot on #452 flagged that install_rootless_docker never configures a
corporate proxy for the user-scoped dockerd (the #244 _configure_docker_proxy
is sudo/system-scoped and the Tier-1 early-return never reaches it), so k3d
pulls of rancher/k3s time out on proxy-only hosts. Deferred to #1221 (the
k3d-on-rootless-socket slice that owns the pulls); leaving a tracked TODO so
the follow-up adds the user-scoped drop-in.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: address Asad review nits on #452 — chmod no-op + misleading comment
- Drop the chmod +x on the rootless installer script: it runs via
`sh "$rootless_script"`, which ignores the exec bit.
- Reword the Tier-1 _install_userspace_tools comment: tools still
sudo-install on Tier 1 (only _persist_tools_on_path is no-sudo);
the comment previously implied otherwise.
The underlying _set_tools_target sudo-crash on no-sudo hosts and the
post-install DOCKER_HOST shell persistence are tracked to #1221.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: resolve user via id -un in install_rootless_docker (Saqlain review, #452)
$USER can be empty in headless / su / cron contexts (a Tier-1 target),
which would break `loginctl enable-linger` and the success line. Resolve
the user once via `id -un` (fallback $USER) and use it for the linger
call, its hint, and the success message. Matches the id-based robustness
DOCKER_HOST already uses. Happy-path bats now mocks `id -un` cleanly.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: add code-quality caller workflow (advisory) (#463)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Add pre-commit hooks (Layer 0, lint-only) (#465)
* Add pre-commit config (Layer 0, lint-only)
Lint-only on purpose: scripts/manifest.sha256 must keep matching the bytes
under scripts/, so no hook may rewrite files.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Document pre-commit setup in README
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs: add Bugbot resolve-and-reply team norm to .cursor/BUGBOT.md (#464)
Part of tracebloc/backend#1308
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* ci: cover scripts/resolve-ingestor-digest.sh in CI shellcheck (#466)
* ci: lint scripts/resolve-ingestor-digest.sh in CI shellcheck (was never linted)
Both CI shellcheck invocations enumerate files explicitly and both omitted
this script. Verified clean against shellcheck --severity=error --shell=bash
0.11.0 before adding. The pre-commit hook from #465 already covers it
locally; this closes the same gap on the CI side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* ci: lint scripts/resolve-ingestor-digest.sh in CI shellcheck (was never linted)
Both CI shellcheck invocations enumerate files explicitly and both omitted
this script. Verified clean against shellcheck --severity=error --shell=bash
0.11.0 before adding. The pre-commit hook from #465 already covers it
locally; this closes the same gap on the CI side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): honest cosign bootstrap download + translated DISM failures (#468) (#469)
The v1.9.7-rc.1 FR killed a healthy install: PS 5.1's progress overlay
throttled the 17 MB pinned-cosign fetch to ~4.5 min of dead silence and
the window read as frozen.
- silence the PS 5.1 progress overlay in Get-WithRetry/Get-Optional
(function-local, auto-reverts) - the classic 10-50x IWR speedup
- run the cosign fetch in a background job with a dim liveness tick
(Wait-JobWithTicks / Get-OptionalWithTicks; cwd pinned per #409,
TLS 1.2 re-applied in the fresh process), expectation lines before,
elapsed + checksum-verified confirmation after
- ASCII-only string literals in both installers: the release asset is
served without a charset so PS 5.1's irm decodes UTF-8 source as
Latin-1 before iex, and BOM-less -File reads are ANSI - literal
em-dashes/ellipses reached customers as mojibake. Locked in by a
tokenizer-based Pester test (which also caught the -Help here-string).
- Enable-OneVirtFeature: translate DISM's raw COMException (feature
package absent on Server SKUs vs enable failure) and stop demanding
a reboot for a feature that never enabled (old code sent Server
users into a reboot->re-run->same-error loop)
Pester: 212 passed / 0 failed locally (pwsh 7.5, Pester 5.7.1).
PSScriptAnalyzer: 0 errors. manifest.sha256 regenerated.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(installer): trust the corporate MITM CA in the k3d nodes (#424) (#453)
* fix(installer): trust the corporate MITM CA in the k3d nodes (#424)
Proxy REACHABILITY reaches the nodes, but on a TLS-inspecting (break-and-
inspect) network the nodes still don't TRUST the corporate CA, so every in-node
containerd pull (rancher/k3s, ghcr.io, tracebloc images) fails x509 — then
masked (helm runs without --wait) into a root-cause-free "an image couldn't be
pulled." Enterprise/hospital archetype, all three OSes.
- Inject the CA at create time: when TRACEBLOC_CA_BUNDLE (or CURL_CA_BUNDLE) is
set, mount the bundle into every k3d node and write a registries.yaml pointing
containerd at it per-registry (docker.io, registry-1.docker.io, ghcr.io), via
the same --config/create path that already carries proxy env. Parity across
scripts/lib/cluster.sh (Linux/macOS) and install-k8s.ps1 (Windows). A CA var
set but unreadable fails loudly instead of silently skipping.
- Name the env var where the user hits the wall: the TLS-interception preflight
hint (both OSes), docs/INSTALL.md, and the PS -Help env-var list.
- CA-aware diagnosis: detect x509 / "certificate signed by unknown authority"
pull events and report a dedicated image_pull_ca state — "the cluster does not
trust your network's TLS-inspection CA" + the exact remedy — instead of the
generic pull error. Mirrored in summary.sh and Print-Summary.
- New check-drift.sh parity check (_drift_ca_trust) so neither installer can drop
the CA wiring for the other's OS.
Tests: +8 cluster.bats, +3 summary.bats, +2 check-drift.bats, +8 Pester.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): CA-trust hardening — no fail-open, bounded events, verify CA readable (Bugbot #424)
Three Bugbot findings on #424:
- _write_k3d_registries_config failed open: on mktemp failure it returned success
with no path, so create still mounted the CA and logged "nodes trust it" but
dropped --registry-config → containerd never got ca_file, x509 pulls still fail
while the operator thinks it's fixed. Now returns non-zero; the caller hard-errors
(CA was supplied, so we refuse to proceed without wiring it in).
- PS Get-NotReadyState `kubectl get events` had no --request-timeout (the bash path
does) — on a wedged/proxy-misrouted API, classification could hang. Added
--request-timeout=5s to match _diagnose_not_ready.
- PS Resolve-CaBundle only checked existence (Test-Path), not readability, so an
unreadable CA passed on Windows but bash (-r) hard-fails. Added an OpenRead probe
so both fail the same way, up front.
Tests: cluster.bats +mktemp-failure + unwritable-registries-hard-error;
install-k8s.Tests.ps1 +unreadable-CA (Unix) + events --request-timeout assertion.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): errexit-safe CA-resolve capture + drift check ignores comments (Bugbot #424 r2)
Two round-2 Bugbot findings:
- Under `set -euo pipefail`, `ca_bundle="$(_resolve_ca_bundle)"; ca_rc=$?` exited on
the rc-2 (unreadable/missing CA) BEFORE ca_rc/error ran — operators got a bare
exit instead of the "can't be read" guidance. Capture with `|| ca_rc=$?` so
errexit doesn't fire and the guidance prints.
- _drift_ca_trust whole-file grep matched tokens in comments (e.g. --registry-config
appears in a comment above the real line), so deleting the functional wiring could
still pass. Strip comment lines first (matches the execute-gate / preflight-host
checks), no grep -q under pipefail.
Tests: cluster.bats +errexit-safe-capture; check-drift.bats +comment-only-token drift.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): TLS-preflight hint names the right var per layer/OS (Bugbot #424 r3)
The hint claimed TRACEBLOC_CA_BUNDLE makes "the host AND the k3d nodes" trust the
CA, but the host connectivity checks use curl_secure / Invoke-WebRequest, which
read CURL_CA_BUNDLE / the system trust store — not TRACEBLOC_CA_BUNDLE (that var
only reaches the nodes via _resolve_ca_bundle). Following the hint literally left
host preflight TLS failures unchanged. Corrected, no behaviour change:
- bash: CURL_CA_BUNDLE fixes these host checks AND the nodes; TRACEBLOC_CA_BUNDLE
is nodes-only; or add the CA to the system trust store.
- Windows: import the CA into the cert store for the host checks (Invoke-WebRequest
uses the store, not an env var); TRACEBLOC_CA_BUNDLE/CURL_CA_BUNDLE cover the nodes.
(Reworded to avoid a bare lowercase `curl` that the curl_secure style guard flags.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): apply CA on cluster REUSE path — warn + recreate guidance (Bugbot #424 r4)
The image_pull_ca remedy said "set the CA and re-run", but CA trust is baked in
only at fresh create; a re-run reuses the existing cluster and never mounts the CA
or passes --registry-config, so the x509 pulls persisted. Mirror the existing
proxy handling (baked-at-create → warn on reuse):
- bash _check_existing_cluster_ca (called from _handle_existing_cluster): warns when
a CA bundle is set but the reused server container lacks the CA mount.
- ps1 New-K3dCluster reuse block: same check via docker inspect mounts.
- both image_pull_ca remedies now say to `k3d cluster delete <name>` first, then
re-run with the CA (CA, like proxy, can't be added to a running cluster).
Tests: cluster.bats +3 (no-CA no-op / CA-but-missing-mount warns / mount-present silent).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(installer): add auth.docker.io to the CA registries config (Bugbot #424 r5)
The registries.yaml ca_file entries covered docker.io / registry-1.docker.io /
ghcr.io, but Docker Hub pulls also TLS-handshake with auth.docker.io for bearer
tokens — so on a break-and-inspect network containerd still rejected the
intercepted cert there even with the CA mounted. #416 already probes auth.docker.io
at preflight; the CA registries list now matches. Added to TB_CA_REGISTRIES and
$TbCaRegistries; registries.yaml test counts 3 -> 4.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#424): _resolve_ca_bundle rejects a directory, not just unreadable paths
A directory of PEMs is readable (-r) but would bind-mount over the single
node ca_file path and containerd can't read it — the silent 'looks applied
but still x509' case. Require a regular file (-f), mirroring the PS
Resolve-CaBundle -PathType Leaf check. Adds a directory-reject bats case.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#424): exact whole-line match for CA mount detection (Bugbot)
_check_existing_cluster_ca used a substring test on docker mount
destinations, so a longer path embedding /etc/ssl/certs/tracebloc-mitm-ca.crt
(e.g. …crt.bak) would be treated as the CA mount and skip the recreate
warning while containerd still x509-fails. Switch to grep -qxF (exact
whole-line), matching the PS anchored regex. Adds a substring-embed test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#424): scope x509 classification to the pull-failure event (Asad)
_diagnose_not_ready / Get-NotReadyState flagged image_pull_ca on ANY x509
event in the namespace, so a stale/unrelated x509 event (e.g. a FailedMount)
could misdirect the user into a needless delete+recreate. Filter events to
the image-pull failure lines (failed to pull / ErrImagePull) before testing
x509, in both bash and PS. Adds an unrelated-x509 test to each side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(chart): perIngestionTables — RFC-0003 D16 enablement knob (backend#1205) (#472)
* feat(chart): perIngestionTables — the RFC-0003 D16 enablement knob (backend#1204/#1205)
values.perIngestionTables (default false, schema-typed) renders
PER_INGESTION_TABLES=1 onto the jobs-manager, which forwards it into
every ingestion Job it spawns (client-runtime companion PR). Flip per
environment, dev first, only once that environment's backend + engine
images + jobs-manager carry the merged D-series. Default installs
render byte-identically (conditional block; unit tests pin both sides).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(values): own banner for perIngestionTables — it is not part of the authz section (review)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* chore: clear house-rules findings (#470)
Fix every finding the shared org checker (tracebloc/.github
scripts/house-rules.sh) reports at develop HEAD: missing curl
timeouts/TLS floors, plus (cli) a missing pipefail. Waivers only where
the finding is a documented false positive. Part of tracebloc/backend#1303.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(install): Tier-1 subuid/subgid gate + prepare-host remediation (#1220) (#458)
* feat(install): Tier-1 subuid/subgid gate + prepare-host remediation
RFC 0001 #1220. Detect the one privileged residue a modern rootless host may
still need — a subordinate UID/GID range + the setuid uidmap helpers — and
either proceed (present), hand off to prepare-host (unprivileged), or perform
one announced touch (sudo available). Never blanket sudo, never an opaque
mid-install crash inside dockerd-rootless-setuptool.sh.
- probe.sh: _probe_subid_ranges (PROBE_SUBID) + _probe_uidmap_helpers
(PROBE_UIDMAP), set in run_host_probes (Linux only), plus audit rows on
the Tier-1 path.
- common.sh: shared pure parsers _subid_has_entry + _next_subid_start, used
by both the probe and the remediation (no duplication).
- setup-linux.sh: _ensure_subid_ranges gate (present / hand-off / one
announced sudo touch) called before install_rootless_docker;
_provision_subid_ranges (idempotent, non-overlapping block, usermod
--add-subuids with file-append fallback, uidmap install) shared by the
installer and run_prepare_host. Folds in slice-1's minimal uidmap check.
- Tests: probe.bats + setup-linux.bats. Manifest regenerated (R8).
Closestracebloc/backend#1220
Part of tracebloc/backend#1177 · Epic tracebloc/backend#1168
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address #458 review — subid gate/probe/hand-off robustness
Bugbot + @saadqbal + a self code-review on client#458, all in slice-2's code:
- id -un everywhere (gate, _provision default, probe): $USER diverges from the
rootless daemon's user under su/cron, which wedged detection/provisioning (#1).
- Re-verify the uidmap helpers are usable (present AND setuid|cap_setuid) after
install, and return non-zero + warn (NOT error/exit) so run_prepare_host stays
best-effort while the installer sudo-path hard-fails via `|| error` (#2 + self-review).
- _idmap_helper_ok (common.sh): accept the setuid bit OR a cap_setuid filecap, so
Arch's `shadow`/pacman path isn't false-rejected (#3).
- Hand-off + run_prepare_host fallback compute a non-overlapping start via
_next_subid_start (honoring TB_SUBUID_FILE/TB_SUBGID_FILE), not hardcoded 100000
(#4 + self-review path-override).
- Hand-off command names the researcher (TB_PREPARE_USER=) — bare prepare-host
provisions nothing, so it would have looped back to the same hand-off (#5).
- Capture `usermod --help` before grepping — pipefail-safe (#6).
bats: id -un mocks, filecaps accept/reject, gate hand-off (names user + computed
start), _provision re-verify best-effort, run_prepare_host best-effort. R8 regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(install): stub the subid gate in the Tier-1 rootless routing test
install_linux's Tier-1 branch now calls _ensure_subid_ranges (slice 2)
before install_rootless_docker; the routing test left it un-stubbed, so the
real gate hit the no-sudo hand-off and error()'d → install_linux returned
non-zero. Stub _ensure_subid_ranges (its own behavior is covered by the
dedicated gate tests) and assert it runs before daemon setup.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(install): scope set -o pipefail to a subshell (bats harness footgun)
Setting `set -o pipefail` in the @test body can leak into bats' own
post-test pipelines and fail the whole run with exit 1 even when every
test reports ok (no 'not ok'). Confine it to a subshell around the call
so the pipefail-safety assertion still holds without touching the harness.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): style guard — no bare curl in the prepare-host hint
The hand-off piped 'curl … | TB_PREPARE_USER=… bash', which breaks
check-style.sh's exemption for the canonical 'curl … | bash' one-liner
(the env var sits between the pipe and bash). Split into an 'export
TB_PREPARE_USER=…' line + the canonical piped one-liner — still names the
researcher, and passes the guard. Verified with scripts/check-style.sh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): fix the #458 red bats + 2 Bugbot bugs (newgidmap cap, write-failure)
Root cause of the "540 ok but exit 1" bats red: the probe.bats uidmap tests set
PATH="$bin" in the test body to hide system helpers, which also hides `rm` — so
bats-core 1.10+ can't run its own per-test cleanup ("rm: command not found") and
fails the whole run even though every test passes. Scope the hermetic PATH to a
subshell so it can't leak into bats' machinery. (Why develop was green + this was
so hard to see: these tests are new in slice 2, and the symptom is a clean pass
list with a non-zero exit.)
Two real Bugbot findings in the slice's own code:
- _idmap_helper_ok checked cap_setuid for BOTH helpers; newgidmap carries
cap_setgid (Arch filecaps) -> false-rejected. Map name->cap; fix the test mock
that masked it + add a wrong-cap regression test.
- _provision_subid_ranges printed success/returned 0 even when the usermod/tee
write failed (callers run it with set -e off) -> installer proceeds with no
range. Guard every write; warn + return 1 on failure. + a test.
Verified: probe.bats + setup-linux.bats EXIT 0 (0 not-ok, 0 rm-not-found) in a
faithful ubuntu 24.04 + bats 1.10 + non-root container. Rebased onto develop.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): 2 more Bugbot findings on #458 (apt hang, false zero-root message)
- _install_uidmap_pkg ran a bare `sudo apt-get install -y uidmap` under the
spinner — no needrestart/DEBIAN_FRONTEND env, no DPkg::Lock::Timeout, no
apt_wait_for_lock — so a headless Tier-1 install can hang on Ubuntu needrestart
or an apt-daily lock (#210 class). Reuse the repo's hardened PM_INSTALL
(populate via setup_pm, which Tier 1 skips) + apt_wait_for_lock.
- install_rootless_docker always printed "no administrator rights were used",
even after _ensure_subid_ranges performed an announced sudo touch on the
root/sudo_nopw path. The gate now sets TB_ROOTLESS_ADMIN_TOUCH and the summary
is honest on both the zero-root and one-admin-touch paths.
Tests: hardened-install assertion (NEEDRESTART_MODE + DPkg::Lock::Timeout) + a
success-message honesty test. Verified EXIT 0 (0 not-ok, 0 rm-errors) in the
faithful ubuntu 24.04 + bats 1.10 + non-root container.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): refresh the package index in _install_uidmap_pkg (Bugbot #458)
Completing the prior apt-hardening: _install_uidmap_pkg populated PM_INSTALL and
waited for the dpkg lock but never ran PM_UPDATE. On the Tier-1 path this is the
first package op, so an empty/stale index can't locate uidmap/shadow and the
install hard-stops. Run $PM_UPDATE (best-effort) first, matching the repo's other
install paths (setup-linux.sh:335/543). Test asserts the index refresh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(installer): silence the PS 5.1 progress throttle in install-k8s.ps1 downloads (#468 follow-up) (#471)
Same class as the bootstrap fix in #469: PS 5.1's progress overlay
throttles Invoke-WebRequest 10-50x and reads like a hang. One
function-local $ProgressPreference in Invoke-WithRetry covers every
fetch scriptblock it drives (dynamic scoping) - winget msixbundle,
Docker Desktop fallback, kubectl, k3d, helm, GPU plugin yaml, and the
version resolvers.
Honest-progress expectation lines (sizes measured today via HEAD):
Docker Desktop ~600 MB, winget ~200 MB, kubectl ~60 MB, k3d ~25 MB,
helm ~20 MB - all cold-path only, silent on warm re-runs.
Pester: 205 passed / 0 failed locally. PSSA: 0 errors.
manifest.sha256 regenerated.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: shujaat hasan <shujaat@tracebloc.io>
* fix(installer): trust the corporate CA in the Docker daemon for k3d runtime image pulls (#474) (#475)
* fix(#474): detect host Docker daemon x509 at cluster-create + document daemon CA trust
k3d pulls its own runtime images (rancher/k3s, k3d-tools, k3d-proxy) with the
HOST Docker daemon, which doesn't use the in-node CA trust from #424. On a
TLS-inspecting network that pull can x509-fail during 'k3d cluster create',
before any node boots — so the post-create diagnosis never classifies it.
- bash: _host_ca_create_hint() detects x509 in the k3d create output and prints
a platform-aware remedy (Linux system trust store vs Docker Desktop VM);
wired into _create_new_cluster's failure path.
- PS: Write-HostCaCreateHint() mirrors it (Windows Trusted Root store), wired
before the generic create failure.
- docs/INSTALL.md: document trusting the CA in the daemon itself (Linux /
Docker Desktop).
- check-drift.sh: enforce both installers keep the host-CA hint (parity).
- Tests: bats (Linux/macOS branches + silent-on-no-x509) + Pester + drift.
Closes#474
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): distro-aware Linux CA remedy + actionable Docker Desktop for Linux (Bugbot)
- Linux native-Docker remedy now covers both Debian/Ubuntu (update-ca-certificates)
and RHEL/Fedora (update-ca-trust), not just the Debian path — the installer
supports RHEL hosts where the Debian commands fail.
- Docker Desktop for Linux now has an actionable step (trust in the system store,
restart Docker Desktop) instead of a dangling reference to a step only printed
on the macOS branch.
- docs/INSTALL.md updated to match. bats Linux test asserts both distro paths.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): cover Colima runtime in the macOS host-CA remedy (Bugbot)
Headless macOS installs use Colima (_install_docker_colima), a Lima VM that
does not read the macOS keychain — so the 'trust it in the keychain + restart
Docker Desktop' remedy was wrong for those hosts. The macOS branch now also
gives the Colima path (add the CA inside the VM via 'colima ssh', then
'colima restart'). docs/INSTALL.md + macOS bats test updated to match.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#474): isolate drift negative fixtures so a missing new token can't mask them (Bugbot)
Adding _host_ca_create_hint / Write-HostCaCreateHint as required _drift_ca_trust
tokens meant the older negative fixtures (missing registry-config, comment-only
registry-config) could pass just because the new token was also absent — so the
comment-strip case no longer uniquely proved comment-stripping still works. Each
negative fixture now carries ALL other required tokens and omits/comments only
the one under test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): surface host-CA hint on the PS create-timeout path too (Bugbot parity)
The PowerShell create-timeout branch exited via Err without calling
Write-HostCaCreateHint (and deleted the k3d logs first), so a TLS-inspected
host pull that logs x509 then hangs to the deadline gave Windows operators a
raw timeout with no certlm.msc CA guidance — while bash runs _host_ca_create_hint
on its timeout fall-through. Capture the full create output before deleting the
logs and call the hint before the timeout Err. Adds a parity regression test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#474): use a herestring in _host_ca_create_hint to survive pipefail (Asad)
printf '%s' "$out" | grep -qiE ... could swallow the hint under set -o pipefail:
grep -q closes the pipe on first match, so for output past the ~64KB pipe buffer
(reachable on the timeout path, which passes the full logs) printf takes SIGPIPE,
the pipeline exits non-zero, and `|| return 0` bails even though x509 matched.
Feed grep via a herestring (no pipe, no SIGPIPE). Adds a >64KB-under-pipefail
regression test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(install): Tier-1 k3d-on-rootless + cgroup delegation + autostart (#1221) (#478)
* feat(install): Tier-1 k3d-on-rootless + cgroup delegation + autostart (#1221)
Slice #1221 (RFC 0001 / #1177): make a rootless Tier-1 cluster actually usable,
all behind the opt-in TB_TIER1_ROOTLESS flag (default off until the spike's §5
host validation). With the flag unset every path below is a no-op and current
behavior is byte-for-byte unchanged.
- Shared _rootless_active predicate (common.sh) so cluster.sh + setup-linux.sh
can't drift on the flag pair.
- create_cluster targets the rootless socket (DOCKER_HOST); ensure_cluster_
autostart gets a user-scope branch (systemctl --user enable + loginctl
enable-linger, never `sudo systemctl enable docker`), and promises
reboot-survival only when BOTH succeed (honesty rule, #375/#458).
- cgroup v2 controller delegation drop-in (Delegate=cpu cpuset io memory pids):
privileged write + daemon-reload on root/sudo, or hand off to prepare-host
with the exact path+content when unprivileged. run_prepare_host writes it too
(system-wide -> covers the researcher).
- Carry-ins from #452/#458: scope-aware _configure_docker_proxy (user scope, no
sudo) so a proxy-only host's rootless daemon can pull rancher/k3s;
_set_tools_target installs user-space on rootless Tier 1 (no sudo-mv crash on
a true no-sudo host); persist DOCKER_HOST to the shell rc for new terminals.
14 new bats tests incl. flag-off regressions; shellcheck --severity=error clean;
manifest.sha256 regenerated (R8).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address Bugbot on #478 — rootless autostart seed + admin-touch msg
- ensure_cluster_autostart: don't seed TB_DOCKER_AUTOSTART from the SYSTEM
docker.service is-enabled check on the rootless path. The cluster runs on the
per-user rootless socket, so a system unit that happens to be enabled would seed
a false reboot promise the rootless branch then can't honestly retract. On
rootless the user-scope enable+linger are now the sole authority (Bugbot medium).
- install_rootless_docker: the TB_ROOTLESS_ADMIN_TOUCH success line no longer
hardcodes "subuid/subgid range" — _ensure_cgroup_delegation can set that flag
too, so it now names "host prerequisites (subuid/subgid range and/or cgroup
delegation)" (Bugbot low).
- Test: rootless + system docker.service enabled + user-enable fails => flag stays
0 (pins the seed-guard). manifest regenerated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(install): address Asad + Saqlain review on #478
- _persist_docker_host: key idempotency off our own marker, not a bare
'DOCKER_HOST=' probe. The old probe also matched a user's own DOCKER_HOST
(remote/TCP), so we silently skipped persisting the rootless socket and new
shells kept hitting the wrong daemon. Now: our own line -> idempotent; a foreign
DOCKER_HOST -> left untouched + a warn to repoint it (Asad #2 + Bugbot #478, Medium).
- ensure_cluster_autostart: reset TB_DOCKER_AUTOSTART=0 in the rootless else-branch
(defensive; the is-enabled seed is already guarded off the rootless path) so the
honesty guarantee is local to the branch (Asad #1).
- install_rootless_docker: success line now reads "one or more one-time admin steps"
so it doesn't undercount when both the subuid and cgroup touches happen (Saqlain #1).
- Test: foreign DOCKER_HOST -> warns, no clobber, no double-write. manifest regen.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(installer): failures surface the real error + log path + support-bundle hint (#423) (#476)
* fix(#423): failures surface the real error + log path + support-bundle hint
Fatal errors printed a generic red line while the actionable detail (k3d/helm
stderr) went only to the transcript, and the log path itself was never shown on
screen. Now:
- Err gains an optional $Detail param; Get-ErrDetailLines (pure, unit-tested)
renders the last ~5 non-empty output lines + the log path + a '-Diagnose'
next-step hint, appended to EVERY fatal error.
- Cluster-create failure passes k3d's stdout/stderr so the real reason (image
pull / proxy / port / WSL) shows on screen — the motivating case.
- Helm repo-add / reconcile / install failures pass helm's output via $Detail
instead of embedding it (no more duplicated log-path text).
- Install log path is announced up front in the banner (was log-only before).
Closes#423
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): force array enumeration of Err detail lines (Bugbot, defensive)
Bugbot flagged that a single-line Get-ErrDetailLines return (no detail + no
LOG_FILE, e.g. a Confirm-Config failure before Start-InstallLog) unwraps to a
scalar string. The foreach statement already iterates a scalar once (verified:
it prints the whole line, not per-character), so the reported char-splitting
does not reproduce -- but wrap the enumeration in @(...) to make that
unambiguous and future-proof. Adds a regression test asserting the single-line
case stays one intact line.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): strip PS 5.1 ErrorRecord chrome from the failure excerpt (Bugbot)
helm failures arrive as `native 2>&1 | Out-String`; on Windows PowerShell 5.1
that wraps stderr in ErrorRecord chrome (the `At <file>:<n> char:<n>` position
line plus the `+ ...` / `+ CategoryInfo` / `+ FullyQualifiedErrorId` block).
Get-ErrDetailLines kept only the last 5 non-empty lines, so the excerpt was all
chrome and the real `Error:` line dropped out -- a regression from the previous
full-message dump. Filter those chrome lines before taking the window so the
actual error survives. Adds a regression test simulating the 5.1 rendering.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): drop duplicate inline log-path hints (Bugbot)
Err now always prints the log path via Get-ErrDetailLines, so the k3d
spawn-failure and create-timeout paths that still Hint "Full log:" right before
Err printed it twice. Remove those inline hints; Err is the single source. Adds
a guard test asserting no inline 'Full log:' hints remain in the installer.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#423): put stderr last in the create-failure Err detail (Asad)
Get-ErrDetailLines keeps the LAST 5 non-empty lines, so with detail ordered
stderr-then-stdout any k3d stdout tail could crowd the real stderr reason
(FATA/x509/port) out of the excerpt. Order it stdout-then-stderr so the stderr
tail survives the window; also matches the Write-HostCaCreateHint order just
above.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Syed Is Saqlain <saqlain.syed007@gmail.com>
Co-authored-by: Syed Saqlain <syedsaqlain@MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Arturo Peroni <arturo@tracebloc.io>
Co-authored-by: shujaat_tracebloc <153823837+shujaatTracebloc@users.noreply.github.com>
Co-authored-by: shujaat hasan <shujaat@tracebloc.io>
Co-authored-by: tracebloc-release-train[bot] <309815517+tracebloc-release-train[bot]@users.noreply.github.com>
* ci(1606): standard-checks and helm-ci call the Makefile instead of copying it
CI parity. Two jobs restated targets the Makefile already declares -- and one
of them had ALREADY DRIFTED, in the direction nobody notices.
LINT. This job spelled out the shellcheck file list inline while the Makefile
kept the same list in SHELLCHECK_FILES. Measured: the Makefile carries 19
entries, this file carried 9. Ten scripts were shellchecked on a contributor
machine and NOT at the merge gate:
gen-manifest.sh check-facts.sh check-style.sh lib/*.sh
tests/check-drift.sh tests/e2e-full-seal.sh tests/e2e-journey.sh
tests/path-persist.sh tests/chart-env-vocabulary.sh
tests/env-vocabulary-agreement.sh
gen-manifest.sh is the installer integrity-manifest generator, so a shell
defect there could not be caught by this gate. `make lint` is green across all
19 on this tree, so arming the full list imports no backlog.
HELM LINT. The values-file loop, `helm lint --strict ./ingestor` and both
vocabulary scripts were verbatim copies. `env-vocabulary-agreement.sh` exists
to prove the four CLIENT_ENV declarations agree with each other (backend#1729
sweep 5) -- a check about "these declarations must not drift" being itself
declared twice is the joke version, and a third vocabulary script added to the
Makefile alone would leave this gate silently not running it.
The apt install of shellcheck stays: it bootstraps the runner, it is not a
duplicated command.
Job names untouched. `Lint` and `Unit tests` are required status checks on
main, matched by name.
Verified: make lint, make helm-lint and make helm-vocab all exit 0 on this
tree; actionlint clean on both files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Asad Iqbal (Saadi) <asad.dsoft@gmail.com>
Co-authored-by: shujaat hasan <shujaathasan@shujaats-MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: shujaat_tracebloc <153823837+shujaatTracebloc@users.noreply.github.com>
Co-authored-by: Arturo Peroni <arturo@tracebloc.io>
Co-authored-by: tracebloc-release-train[bot] <309815517+tracebloc-release-train[bot]@users.noreply.github.com>
Co-authored-by: Syed Is Saqlain <saqlain.syed007@gmail.com>
Co-authored-by: Syed Saqlain <syedsaqlain@MacBook-Pro.local>
Co-authored-by: shujaat hasan <shujaat@tracebloc.io>
shujaatTracebloc added a commit that referenced this pull request Aug 20, 2026
…erdicts (#772) (#773)
client#766 gave the two installers a shared CONTRACT (envelope_contract.json) and
that fixed arithmetic agreement. It did not give them shared CONTROL FLOW, and
every divergence since has landed in that gap -- five across backend#2220:
1. node ordering (memory,cpu) in bash vs (cpu,memory) in the CLI/ps1, so the
two anchored on DIFFERENT nodes on a heterogeneous cluster
2. [math]::Max(0,..) binding Int32 on ps1, swallowed by a bare catch{} --
machine sizing was silently DEAD on Windows and nothing failed
3. ps1 coerced unparseable allocatable to 0 and RANKED the node; bash skipped it
4. split provenance lookup: on a failed read ps1 said `installer` (invites a
ladder to overrule a human) while bash said `unknown` (permanently strands
an installer-sized edge). Opposite directions, so neither twin's behavior
told you anything about the other's.
5. bash INFERRED "too small" from a second `kubectl get nodes -o name` probe,
so unparseable or not-yet-Ready nodes tripped a warning about a machine
never measured; ps1 gated on a parsed node.
Each was caught one at a time, by review or Bugbot, after the code shipped. That
worked, but it is luck-shaped: #2 had been live and silently disabling machine
sizing on Windows, and surfaced only because someone happened to write a
full-matrix replay. Track A adds more logic to this same path, so the divergence
surface is about to grow.
The golden vectors cover the arithmetic. What they cannot cover is behavior on
everything that is NOT a clean measurement -- a node that will not parse, a
values read that fails, a remainder too small to request. Those are exactly the
states all five findings lived in, and each twin decided them alone.
So: fixtures/installer_parity.json, 18 cluster states, each declaring four
verdicts (size, provenance, undersized, unschedulable). installer-parity.bats
drives _resolve_training_size; installer-parity.Tests.ps1 drives
Get-TrainingResources/Get-TrainingProvenance. One table, two readers -- a row
added to the JSON forces BOTH languages to answer it.
Both suites stub at the same boundary (the two external commands, kubectl and
helm) rather than stubbing the installers' own helpers, so the two sides cannot
drift in what they actually exercise -- which would reintroduce the very problem
this closes. PowerShell reads the JSON directly; bash reads a generated table,
because jq is not a prerequisite -- the same split envelope_vectors.bash already
uses.
PROVEN to catch the class, in both directions, rather than assumed:
* reverted divergence 3 into ps1 only -> ps1 failed on exactly
one-unparseable-one-valid and all-nodes-unparseable; bash stayed green
* reverted divergence 1 into bash only -> bash failed on exactly
heterogeneous-incomparable and node-order-reversed; ps1 stayed green
Both installers were then restored; this commit does not touch either.
The fixture also declares excluded_from_parity WITH reasons, and Pester asserts
those reasons exist. Parity that quietly skips the awkward states is worse than
no parity, because it reads as coverage. The one exclusion is "kubectl binary
absent": bash gates on its own `has` helper while ps1 infers absence from
$LASTEXITCODE, so comparing them would compare two different questions.
No CI wiring needed: installer-tests.yaml already runs `Run.Path = scripts/tests`
and standard-checks runs `bats scripts/tests/*.bats`, so both halves are
discovered, and gen-installer-parity.sh --check rides in as bats test 1.
Verified: parity 4/4 both sides, all 18 rows; Pester 761 passed / 0 failed (+3,
and the two .Tests.ps1 files coexist); bats 1293 ok with the same 2 pre-existing
failures clean develop has (assess.bats:40, install-bootstrap.bats:88 --
confirmed by stashing); bats-hygiene 18/18; make lint clean over 48 files.
Purely additive -- 5 new files, no installer touched, no manifest change (dev
tooling and tests are not on the signed file surface).
Closes#772
Refs: backend#2220, client#766, client#768, backend#664
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saadqbal@divyasinghds