fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct - #515

Closed
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica
Closed

fix(x402): verifier replicas: 2 → 1 to keep metric GC correct#515
bussyjd wants to merge 1 commit into
feat/x402-marketplace-metricsfrom
fix/verifier-single-replica

Conversation

@bussyjd

Copy link
Copy Markdown
Contributor

Why

0fbb99a (this PR's parent) shipped pruneSeriesNotIn to GC verifier metric series when offers are deleted. The GC is per-pod (each pod runs its own informer + its own registry). With replicas: 2 + ServiceMonitor scraping Endpoints (round-robin), Prometheus sees inconsistent series — deleted offers come back every other scrape.

Before

 ServiceMonitor scrape (round-robin Endpoints)
│
┌───────┴───────┐
▼ ▼
verifier-pod-A verifier-pod-B
metric registry metric registry
GC ran on GC ran on
Reload event ✓ Reload event ?
│ │
▼ ▼
clean series stale series for
for deleted X deleted offer X
│
Prometheus sees flip-flop
→ alerts fire/silence on dead labels
→ dashboards show resurrected offers

After

 ServiceMonitor scrape → only verifier-pod-A
│
▼
single registry
single GC source
│
▼
deterministic series state
alerts trustworthy
dashboards consistent

What changed

  • x402.yaml verifier Deployment: replicas: 2 → 1
  • Removed verifier PDB (was minAvailable: 1 at replicas: 1 — blocks voluntary drains on the only pod, useless on single-node k3d)
  • Added explanatory comment so this isn't re-bumped without thought

Future HA path

If/when this stack runs multi-node and wants 2+ verifier replicas, the correct pattern is ServiceMonitor → PodMonitor (each pod scraped with a pod label) + recording rules using sum without(pod). Not now; correctness > theoretical HA on a single-node k3d.

Test plan

  • go build ./... clean
  • go test ./internal/x402/... green
  • Manual on next stack up: scrape /metrics from the verifier, delete a ServiceOffer, scrape again — series should disappear and STAY disappeared

Stacks on

PR #513 (introduces pruneSeriesNotIn). Will rebase onto main after #513 merges.

Commit 0fbb99a (fix(x402): GC verifier metric series for deleted offers)
added pruneSeriesNotIn to Verifier.load. Each verifier pod runs its
own informer + its own metric registry, so the GC is per-pod. With
replicas: 2 + ServiceMonitor (round-robin scrape over Endpoints),
Prometheus sees:
* one pod's registry on scrape N (pruned correctly),
* the other pod's on scrape N+1 (may still hold a deleted offer's
series until that pod's informer also sees the delete).
Result: deleted offers' last_payment_success_seconds gauge and
charged_requests_total counters reappear every other scrape, polluting
dashboards and creating spurious alert state.
Cheapest correct fix is replicas: 1. The verifier is on the request
path but single-node k3d gains no HA from 2 replicas. Drop the
PodDisruptionBudget too — minAvailable:1 at replicas:1 just blocks
voluntary drains on the only pod, useless on k3d.
If/when the stack ever runs multi-node and HA replicas are wanted,
the right pattern is ServiceMonitor → PodMonitor with a `pod` label
and recording rules using `sum without(pod)`. That's a future change;
right now correctness > theoretical HA.
@bussyjd

Copy link
Copy Markdown
ContributorAuthor

Superseded by bundle PR #536 — closing in favor of the consolidated merge target. Original branch and history preserved.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@bussyjd