fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir - #522

Closed
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc
Closed

fix(x402-buyer): persist consumed-nonce state to PVC instead of emptyDir#522
bussyjd wants to merge 1 commit into
mainfrom
fix/x402-buyer-state-pvc

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

Today /state on the x402-buyer sidecar is emptyDir: {}. Pod restart loses consumed.json → buyer treats every pre-signed auth as fresh → tries to spend already-consumed nonces → facilitator 400 cascade → buyer pool exhausted → 503 until manual buy.py process --all.

Before

 litellm pod restart (rollout, OOM, anything)
│
▼
emptyDir /state wiped auth-pool ConfigMap
│ (controller-managed)
│ │
│ ▼
└──────────────────► buyer reads pool, sees N unused
│
▼
spends auth #1
│
▼
facilitator: "nonce already used" 400
│
▼
buyer 402 → caller retry → buyer spends auth #2
│
▼
facilitator 400 → ... → pool exhausted → 503
│
▼
manual `buy.py process --all` to reseed

After

 litellm pod restart
│
▼
PVC remounted (local-path, RWO, 50Mi) auth-pool ConfigMap
│ │
▼ ▼
consumed.json INTACT buyer reads pool
│ │
└──────────────────► buyer skips consumed entries
│
▼
spends next unused auth
│
▼
facilitator OK → settlement proceeds

What changed

  • llm.yaml — new PVC x402-buyer-state (50Mi, local-path, RWO); volume entry on litellm Deployment swaps from emptyDir: {}persistentVolumeClaim
  • llm.yaml Deployment strategy → Recreate (RWO PVC requires no surge)
  • internal/embed/embed_buyer_state_test.go — regression test pinning the PVC + Recreate invariants

What this does NOT solve

  • Multi-replica litellm (RWO can't be shared; would need StatefulSet OR RWX storage class)
  • Hard node loss (local-path is node-local; on k3d single-node, full blast radius anyway)

PSS compatibility

PR #12 (Restricted PSS sweep) will need to verify the buyer can read/write the PVC mount under runAsUser: 65532. Local-path creates files at 0777 by default which is permissive. Cross-PR coordination noted.

Test plan

  • go build ./... clean
  • go test ./internal/embed/... ./internal/x402/buyer/... green
  • Manual on next stack up: deploy → buy 5 auths → spend 2 → kubectl rollout restart deploy/litellm -n llm → spend 3rd auth → facilitator settles cleanly (no 400)

Today the x402-buyer sidecar's /state directory is an emptyDir. When
the litellm pod restarts (rollout, OOM, node drain), consumed.json is
gone. The pre-signed auth pool reloads from the ConfigMap the
controller manages, and the buyer treats every auth as unconsumed —
attempting to spend nonces that the facilitator already marked used.
Cascade: facilitator returns 400 "nonce already used" -> buyer 402
back to LiteLLM -> caller retry -> same 400 -> eventually buyer pool
exhausted -> 503 until manual `buy.py process --all` reseeds.
Fix: convert /state to a PVC backed by local-path-provisioner (the
storage class already deployed via base/templates/local-path.yaml).
50Mi request; consumed.json is tiny but room left for log growth.
Deployment strategy switched to Recreate because a RWO PVC can't
be co-mounted during a RollingUpdate surge. Litellm is replicas: 1
so this just means rollouts have a ~5s gap instead of an overlap —
acceptable.
What this does NOT solve:
- Multi-replica litellm. RWO PVC works only for replicas: 1; would
need RWX (which local-path doesn't support — needs NFS/Longhorn)
or per-replica state via StatefulSet. Out of scope; litellm has
no current scaling need.
- Hard node loss. local-path PVCs are node-local; if the k3d node
is destroyed, state is gone (along with the rest of the cluster).
For local-only operator that's the expected blast radius.
PSS compatibility note: the PVC mount works under PSS Restricted as
long as the buyer container runs with appropriate fsGroup. PR #12
(Restricted PSS sweep) handles that separately and will verify mount
permissions when it lands.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

@bussyjdbussyjd closed this May 24, 2026
@OisinKyne
OisinKyne deleted the fix/x402-buyer-state-pvc branch July 1, 2026 12:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd