Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) by ducnmm · Pull Request #352 · MystenLabs/MemWal · GitHub
Skip to content

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351) - #352

Closed
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price
Closed

fix(relayer): recover Walrus register destroy_zero from stale WAL price (#351)#352
ducnmm wants to merge 2 commits into
devfrom
fix/walrus-register-destroy-zero-stale-price

Conversation

@ducnmm

Copy link
Copy Markdown
Collaborator

Summary

Fixes#351 — mainnet writes were failing 100% at the register_sponsor phase with an Enoki dry_run_failedMoveAbort in 0x2::coin::destroy_zero (ENonZero).

Not an Enoki fault.@mysten/walrus's #withWal helper pre-funds an exact WAL payment (storageUnits × price × epochs) from the client's cachedsystemState price, then asserts the coin is empty via coin::destroy_zero. When mainnet storage/write price drifts down between the cached read and execution, the contract deducts less WAL than we split off and the leftover trips destroy_zero. Enoki only surfaces it during its budget dry-run.

Evidence (prod, 2026-07-02)

  • On-chain storage_price_per_unit_size moved 71464 → 70922 within ~15 min.
  • After a relayer restart, 6 writes succeeded then the same sidecar client flipped to 100% destroy_zero at the next price move — a time/price transition, not per-user/blob/wallet.
  • classification=permanent retryable=false, so every write was Dead-marked with no retry.

Why it didn't self-heal (two gaps)

  1. The sidecar auto-refresh only fired on isMoveAbortBalanceSplit (matches balance+split); this message says destroy_zero, so refreshWalrusClient() never ran.
  2. The Rust worker swept it into the MoveAbort → Permanent catch → Apalis never retried.

A newer SDK does not fix it: @mysten/walrus 1.2.x has no cost/#withWal/destroy_zero change (installed 1.1.7, latest 1.2.3).

Fix

  • enoki.ts — add isMoveAbortWalDestroyZero detector.
  • walrus-upload.tsrefreshWalrusClient() on the abort so the retry rebuilds against the live price.
  • jobs.rs — classify the register destroy_zero abort as Transient (retryable), disjoint from the balance::split gas-budget path.
  • config.ts — drop WALRUS_CLIENT_MAX_AGE_MS default 30m → 60s so the cached price tracks a live-drifting mainnet price (still env-overridable).

Testing

  • New TS detector suite (7 cases) + Rust classification test using the verbatim prod error.
  • Full suites green: 85 TS, 40 jobs Rust tests.

Deployment

Already hot-deployed to the prod relayer (Railway, deployment 26742fd0). Verified live: clean sidecar boot, writes succeeding, and refreshed reason=max_age lines confirm the new 60s max-age is active. This PR lands the same change on dev.

Follow-up worth filing upstream to Mysten: the exact-change #withWal + coin::destroy_zero pattern is inherently racy against a moving on-chain price — it should tolerate leftover (return change) or add a small buffer.

…ce (#351)
Mainnet writes were failing 100% at the register_sponsor phase with an Enoki
dry_run_failed -> MoveAbort in 0x2::coin::destroy_zero (ENonZero). Root cause is
not Enoki: @mysten/walrus '#withWal' pre-funds an *exact* WAL payment computed
from the client's cached systemState price, then asserts the coin is empty via
coin::destroy_zero. When mainnet storage/write price drifts down between the
cached read and execution, the contract deducts less WAL than we split off and
the leftover trips destroy_zero. Verified on prod: on-chain price moved
71464->70922 within ~15 min, and a fresh relayer boot served 6 writes before the
same client flipped to 100% destroy_zero at the next price move.
Two gaps let this fail hard instead of self-healing:
- The sidecar's auto-refresh only fired on isMoveAbortBalanceSplit ('balance'
+ 'split'); this message says destroy_zero, so refreshWalrusClient never ran.
- The Rust worker swept it into the MoveAbort -> Permanent catch, so Apalis
Dead-marked every write with no retry.
Fix:
- enoki.ts: add isMoveAbortWalDestroyZero detector.
- walrus-upload.ts: refreshWalrusClient() on the destroy_zero abort so the retry
rebuilds against the live price.
- jobs.rs: classify the register destroy_zero abort as Transient (retryable),
disjoint from the balance::split gas-budget path.
- config.ts: drop WALRUS_CLIENT_MAX_AGE_MS default 30m -> 60s so the cached
price tracks a live-drifting mainnet price (still env-overridable).
Tests: new TS detector suite (7) + Rust classification test using the verbatim
prod error. Full suites green (85 TS, 40 jobs).
The sidecar hardcoded getJsonRpcFullnodeUrl(mainnet) and ignored the SUI_RPC_URL
env, so it could not be pointed at a dedicated RPC when the public fullnode pool
degrades. Observed today: the public pool served stale reads (a blob certified
at 00:49:14 read back as 'does not exist' at 00:49:49, 35s later), failing
uploads at get_blob/certify/metadata. Honor an explicit SUI_RPC_URL override so
ops can switch to a healthy RPC via env; fall back to the network default.
@DrVelvetFog

Copy link
Copy Markdown

Reviewed this against the SDK internals — diagnosis and fix both look right. The two-gaps writeup nails why it didn't self-heal; I'd independently landed on the same destroy_zerobalance::split distinction, and keeping the two detectors disjoint is the load-bearing call — a split ENotEnough is genuinely ambiguous with gas-pool exhaustion, so you don't want a starved gas wallet reclassified as recoverable stale-price. Register destroy_zero → Transient plus the 30m → 60s max-age both read as correct for tracking a live-drifting price.

On the upstream follow-up you flagged — it's already in flight, and it's exactly what you described:

So once that lands and you bump the SDK, the recovery here becomes the defense-in-depth layer: the buffer means the down-drift destroy_zero mostly stops firing in the first place, and your refresh-and-retry still catches anything that drifts past it. Good to have both sides covered.

LGTM.

@harrymove-ctrl
harrymove-ctrl self-requested a review July 9, 2026 05:16
@ducnmm

Copy link
Copy Markdown
CollaboratorAuthor

Superseded by #389 — both commits of this branch are included there (the deletion feature was stacked on this fix).

@ducnmmducnmm closed this Jul 11, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mainnet writes fail on both prod relayers: Enoki dry_run_failed (MoveAbort balance::destroy_zero in sponsored upload)

3 participants

@ducnmm@DrVelvetFog@harrymove-ctrl