fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321) - #707

Closed
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime
Closed

fix(llm): zero-downtime LiteLLM config operations — hot-path Reloader stance + x402-buyer split (#321)#707
bussyjd wants to merge 3 commits into
mainfrom
fix/litellm-zero-downtime

Conversation

@bussyjd

@bussyjdbussyjd commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

LiteLLM zero-downtime config operations (#321 phases 1+2)

Problem

The hot /model/new path shipped in #320 was being defeated by two later changes:

  1. 2eed34b added a Reloader annotation on litellm-config. Every model add/remove/prefer and every first-time purchase patches that ConfigMap (persistence-first ordering), so Reloader rolled the pod seconds after the hot-add succeeded.
  2. 5c9a879 moved the buyer's consumed-auth state to an RWO PVC inside the litellm pod, forcing replicas: 1 + strategy: Recreate — every rollout became a full inference gap.

Net effect: every config operation took the inference gateway down, which is exactly what #321 set out to fix (and what the latest issue comment asks about).

Phase 1 — stop self-inflicted restarts

  • Reloader now watches litellm-secrets only. litellm-config changes are hot-applied via /model/new + /model/delete by both the CLI and the serviceoffer-controller; the ConfigMap is persistence for the next pod start.
  • obol model prefer no longer restarts LiteLLM (order is an obol convention read from the ConfigMap by Rank(); LiteLLM's router ignores it).
  • RestartLiteLLM fails loudly on rollout timeout instead of warning and reporting success (LiteLLM reliability: hot-reload, zero-downtime restarts, and single-replica fragility #321 item 3).
  • Drift safety net replacing the Reloader annotation: obol model status now compares the ConfigMap model_list against the live router /v1/models and reports missing/extra entries (internal/model/drift.go), so a silently-failed hot call is visible instead of masked by a restart.
  • ReconcileRecorded's ConfigMap-only branch restarts explicitly now that Reloader won't.

Phase 2 — make the remaining restarts gapless

  • x402-buyer is split out of the litellm pod into its own Deployment (1 replica, Recreate, keeps the RWO PVC and single-writer auth semantics) + ClusterIP Service x402-buyer.llm.svc:8402.
  • litellm is now statelessRollingUpdate maxSurge: 1, maxUnavailable: 0. Secret rotation (via Reloader) and image bumps surge a new pod to Ready before the old one terminates — zero inference gap. Still replicas: 1 steady-state (no extra laptop footprint).
  • Paid-route api_basehttp://x402-buyer.llm.svc.cluster.local:8402/v1 (chart wildcard + controller-written entries).
  • Upgrade path: the controller migrates legacy 127.0.0.1:8402 entries at startup and on per-purchase reconcile (ConfigMap rewrite + hot delete/re-add). Those entries are dead upstreams on upgraded clusters until migrated, so the delete/re-add window is not a regression.
  • Controller buyer probes (/admin/reload, /admin/remove, /status) and agent buy.py target app=x402-buyer pods; flow scripts reach the buyer via its Service.
  • PodMonitor keeps its historical name (litellm-x402-buyer) but targets the buyer pods.

Deliberately deferred

  • NetworkPolicy on :8402 — agent namespaces (hermes-obol-agent, openclaw-<id>, agent CRs) share no common label and buy.py legitimately reads /status from them. The Service changes addressing, not exposure (the pod IP was already cluster-reachable). Follow-up: label agent namespaces, then restrict.
  • Hot API-key rotation (fork /model/update with in-memory key) — unnecessary now that key rotation rolls gaplessly; would only remove the surge itself.
  • Buyer HA (auth-pool sharding) — buyer rollouts are rare (image bumps only; CM changes hot-reload) and brief.

Review boundaries / invariants

  • x402-buyer must have exactly one writer of consumed-auth state: buyer Deployment is replicas: 1 + Recreate, PVC unchanged.
  • litellm Deployment must mount no PVC and carry no buyer container (pinned by TestBuyerStatePVC).
  • Reloader annotation must not include litellm-config (pinned by TestLLMTemplate_IncludesPaidRouteAndBuyerSidecar).
  • Controller must never fall back to a pod restart on hot-add failure (unchanged behavior, addLiteLLMModelEntry).
  • /v1 suffix on buyer api_base is load-bearing (CLAUDE.md pitfall 6; pinned by TestBuyerAPIBase).
  • No changes to payment verification, settlement, RBAC, or route exposure.

Tests

  • internal/model/drift_test.go — table-driven DiffRouterModels (wildcards, missing, extra).
  • internal/serviceoffercontroller/purchase_migration_test.go — legacy api_base migration (startup + per-purchase + no-op + missing-CM), buyerAPIBase.
  • Updated structural pins: internal/stack/stack_test.go (annotation stance, service api_base), internal/embed/embed_buyer_state_test.go (buyer Deployment owns PVC, litellm stateless + RollingUpdate params).
  • Full go test ./... green.

Local validation (fresh k3d cluster, OBOL_DEVELOPMENT=true, dev-built CLI + controller image)

An in-cluster probe hit litellm.llm.svc:4000/health/readiness twice a second for the whole session: 1253 samples, 0 failures across all of the following.

  1. Topology: litellm pod is single-container, RollingUpdate {maxSurge:1, maxUnavailable:0}, annotations carry secret.reloader.stakater.com/reload only; x402-buyer runs as its own 1-replica Recreate Deployment with the PVC, plus ClusterIP Service.
  2. obol model prefer qwen3.5 → ConfigMap reordered, same pod, restartCount 0.
  3. obol model remove nomic-embed-text → hot-delete, same pod; obol model status drift check stayed clean (live router really dropped it).
  4. Drift net: appending a bogus entry straight to the CM (bypassing the hot API) made obol model status print missing from router: drift-test-ghost with the fix hint; restoring the CM cleared it.
  5. Secret rotation: patching litellm-secrets triggered a Reloader surge rollout — old and new pods briefly coexisted, buyer pod untouched, zero probe failures through the replacement.
  6. Buyer via Service: /healthzok, /status{}.
  7. Upgrade migration, live: injected a paid/zdt-test entry with the legacy 127.0.0.1:8402/v1 api_base, restarted the controller → log buyer-migrate: rewrote 1 legacy buyer api_base entries in llm/litellm-config, CM rewritten to the Service URL, entry hot-added.
  8. End-to-end path: calling paid/zdt-test through LiteLLM returned the buyer's own application error (no purchased upstream mapped for requested model) — LiteLLM → x402-buyer Service connectivity confirmed; only an actual purchase is absent.
  9. Cleanup verified: obol model remove paid/zdt-test (live + config), final drift check clean.

Not exercised live: a real paid buy (needs funded wallet + seller; covered by release-gate flows 06/08/11/13/14, which this PR updates for the Service address).

bussyjd added 3 commits July 6, 2026 16:14
Phase 1 of the zero-downtime work for #321:
- obol model prefer no longer rolls the LiteLLM deployment. model_list
order is an obol convention (Rank reads the ConfigMap); LiteLLM's
router does not use it, so the restart was pure downtime.
- RestartLiteLLM now fails loudly when the rollout does not converge
within 90s instead of warning and reporting success (#321 item 3).
- ReconcileRecorded's ConfigMap-only branch restarts explicitly now
that Reloader no longer watches litellm-config.
- New drift safety net: obol model status compares the ConfigMap
model_list against the live router /v1/models (CheckRouterDrift /
DiffRouterModels) and reports missing/extra entries, replacing the
Reloader annotation as the guard against silently-failed hot calls.
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
…rollouts
Phase 2 of #321. The buyer's RWO consumed-auth PVC forced the litellm
Deployment onto replicas:1 + Recreate, turning every rollout (Reloader
secret rotation, image bump, config reload) into a full inference gap —
and the Reloader annotation on litellm-config triggered such a rollout
on every model add/remove/prefer and first-time purchase, defeating the
hot /model/new path shipped in #320.
- x402-buyer becomes its own Deployment (1 replica, Recreate, keeps the
PVC and single-writer auth semantics) + ClusterIP Service. Buyer CM
changes still hot-reload via /admin/reload; only buyer image bumps
briefly gap paid/* routes.
- litellm is now stateless: RollingUpdate maxSurge:1 maxUnavailable:0 —
a new pod must be Ready before the old one terminates. Reloader
watches litellm-secrets only; litellm-config changes are hot-applied.
- Paid-route api_base moves to http://x402-buyer.llm.svc.cluster.local:8402/v1.
The controller migrates legacy 127.0.0.1:8402 entries at startup and
on per-purchase reconcile (upgraded-in-place clusters), hot-syncing
the live router.
- Controller buyer probes (reload/remove/status) and buy.py target the
x402-buyer pods; flow scripts reach the buyer via its Service.
- PodMonitor keeps its historical name but targets the buyer pods.
NetworkPolicy for :8402 is deliberately deferred: agent namespaces have
no common label and buy.py legitimately reads /status from them; the
Service changes addressing, not exposure (pod IP was already
cluster-reachable).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
CLAUDE.md and obol-stack-dev references: buyer is a standalone
Deployment + Service (port-forward svc/x402-buyer), paid routes point
at the buyer Service with the mandatory /v1 suffix, and Reloader
watches litellm-secrets only — litellm-config changes are hot-applied
with drift surfaced by obol model status (#321).
Claude-Session: https://claude.ai/code/session_01YLuXwN1A4tG3xmEMKvAzeD
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Closing as superseded by integration/v0.13.0-rc0 @ a29c826 (rolled up in #716 / release v0.13.0-rc0). All commits of this PR are contained in the RC tip (verified by ancestry/patch-id audit).

@bussyjdbussyjd closed this Jul 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd