fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(cli): make cluster-running probe authoritative via kubectl, not file presence - #537

Merged
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe
May 24, 2026
Merged

fix(cli): make cluster-running probe authoritative via kubectl, not file presence#537
bussyjd merged 1 commit into
mainfrom
fix/cli-cluster-running-probe

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Repro

  1. k3d cluster stop obol-stack-<name>
  2. k3d cluster start obol-stack-<name>
  3. kubectl get pods -A (against $OBOL_CONFIG_DIR/kubeconfig.yaml) succeeds
  4. obol sell http <name> --upstream foo --port 80 --namespace default --per-request 0.001 --chain base-sepolia --wallet 0x... fails with:
✗ cluster appears to be stopped — run 'obol stack up' before creating an HTTP service offer

Same false-positive surfaces on every obol sell *, obol network *, obol model *, obol agent *, and most other gated subcommands (every caller of kubectl.EnsureCluster).

What the old probe checked + why it lied

internal/kubectl/kubectl.EnsureCluster only did:

if_, err:=os.Stat(filepath.Join(cfg.ConfigDir, "kubeconfig.yaml")); os.IsNotExist(err) {
returnerrors.New("cluster not running. Run 'obol stack up' first")
}
returnnil

i.e. it asserted "cluster is up" purely from the kubeconfig file existing. The actual cluster appears to be stopped message comes from wrapClusterDown, which sees connection refused on the first kubectl exec the subcommand attempts (e.g. kubectl apply for the ServiceOffer).

After k3d cluster stop && k3d cluster start the kubeconfig file is still on disk but its embedded server: https://0.0.0.0:<port> line points at the previous k3d API port. New port → connection refused → false cluster appears to be stopped. This is pitfall #1 in the project CLAUDE.md ("Kubeconfig port drift — k3d API port can change between restarts").

What the new probe checks + why it's authoritative

EnsureCluster now:

  1. Confirms the kubeconfig file exists (unchanged early-out for never-up clusters).
  2. Actively probes the Kubernetes API server via kubectl version --request-timeout=3s -o json against $OBOL_CONFIG_DIR/kubeconfig.yaml. The same path every downstream kubectl exec will take — if this succeeds, every other call against the same kubeconfig will too.
  3. On a connection refused / no route to host / Unable to connect to the server failure (the existing wrapClusterDown signature), runs one best-effort k3d kubeconfig write <cluster> -o <kubeconfig> --overwrite to recover from the port drift case, then re-probes.
  4. Only if the post-refresh probe still fails does it return ErrClusterDown.
  5. Non-cluster-down failures (e.g. missing kubectl binary) pass through verbatim instead of being masked as "cluster appears to be stopped".

The refresh helper declines silently if prerequisites are missing (no k3d binary, no .stack-id, or .stack-backend is set to a non-k3d backend like k3s) — so the change is safe for the k3s backend and for early-init states.

The probe and refresh are swappable via two package-level vars (probeAPIServerFn, refreshKubeconfigFn), so the recovery branches are fully unit-tested without a live cluster.

Test plan

  • go build ./... clean
  • go test ./internal/kubectl/... -count=1 green (new table covers: probe success, port-drift recovery, refresh skipped, refresh ran but probe still failing, non-cluster-down passthrough, refresh prerequisite checks)
  • go test ./cmd/obol/... -count=1 green (no existing test regressions; tests that seed a stub kubeconfig do not hit EnsureCluster paths that depended on the old no-op behavior)
  • go test ./internal/... -count=1 green (one pre-existing failure in internal/stack/TestWarnIfNoChatModel_EmitsWarnWhenNoModels reproduces on main unmodified — unrelated to this change)

Manual repro test

Live verification on a real cluster (recommend executing before merge):

obol stack up
k3d cluster stop obol-stack-<id>
k3d cluster start obol-stack-<id>
obol sell http demo --upstream ollama --port 11434 --namespace llm \
--per-request 0.001 --chain base-sepolia --wallet 0xYOUR_WALLET

Expected: succeeds after the kubeconfig auto-refresh, no cluster appears to be stopped message. (Old behavior: false positive.)

To confirm the kubeconfig was actually refreshed, diff $OBOL_CONFIG_DIR/kubeconfig.yaml before/after the failing sequence — the server: URL port will have updated.

…ocker labels
The previous EnsureCluster only stat'd kubeconfig.yaml on disk. After
`k3d cluster stop && k3d cluster start` the kubeconfig still exists but
the k3d API port can drift (pitfall #1), so the next kubectl exec gets
"connection refused" and wrapClusterDown wrongly tells the user the
cluster is stopped — even though `kubectl get pods -A` against a
refreshed kubeconfig succeeds.
Replace the file-presence check with an active probe of the K8s API
server (`kubectl version --request-timeout=3s`). On a cluster-down
signature, attempt one best-effort `k3d kubeconfig write --overwrite`
and re-probe before giving up. Non-cluster-down probe failures (e.g.
missing kubectl binary) pass through verbatim instead of being masked
by the misleading "cluster appears to be stopped" hint.
Probe and refresh are swappable via package-level vars so the recovery
branches are fully unit-testable without a live cluster.
@bussyjd
bussyjd merged commit 5bfd042 into mainMay 24, 2026
7 checks passed
@OisinKyne
OisinKyne deleted the fix/cli-cluster-running-probe branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd