Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run - #2689

Merged
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable
Jun 29, 2026
Merged

Backport #2445: Redrive on transient workflow-server transport failures instead of failing the run#2689
pranaygp merged 1 commit into
stablefrom
backport/pr-2445-to-stable

Conversation

@github-actions

Copy link
Copy Markdown
Contributor

Automated backport of #2445 to stable (backport job run).

AI recommendation: This is a self-contained bug fix to existing functionality — transient workflow-server transport failures were being misclassified as USER_ERROR and permanently failing runs, and this redrives them via the queue instead. The modified files (classify-error.ts, runtime.ts, start.ts, queue.ts, http-client.ts, utils.ts, events-v4.ts) all exist on stable, and the change doesn't depend on main-only APIs.

Merge conflicts were resolved by AI (opencode with anthropic/claude-opus-4.8). Please review the conflict resolution carefully before merging.

…iling the run (#2445)
* Redrive on transient workflow-server transport failures instead of failing the run
A firewall in front of workflow-server shedding load with sustained 429/503
makes undici's shared RetryAgent exhaust its retries and throw
UND_ERR_REQ_RETRY. That raw error was rethrown unwrapped, so it was
classified as USER_ERROR and the replay terminal branch wrote run_failed —
permanently failing a run on a transient blip (or, in an outage, falling back
to the ~5min queue visibility-timeout redrive).
- world-vercel: map exhausted-retry / socket / connect / DNS / timeout
failures to a typed WorkflowWorldError (code TRANSPORT/TIMEOUT) by walking
the fetch() cause chain.
- core: add isRetryableWorldError (429 / 5xx / TRANSPORT / TIMEOUT) and
rethrow such errors from the replay terminal branch so the queue redrives
quickly (1s->60s backoff) instead of failing the run. Reuse it in start()
and step_started handling.
- world-vercel: surface the Vercel firewall x-vercel-mitigated
(challenge/deny) header alongside x-vercel-id in error diagnostics and logs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: back off + cap on every retry path; fix mock; refine scope
Revises the transport-error handling per PR review (VaguelySerious,
karthikscale3).
Blocking fix — step_started no longer self-enqueues a throttled defer for
transient world errors. Returning `{ type: 'throttled', timeoutSeconds: 1 }`
acked the delivery and enqueued a fresh message, resetting the delivery count
so the path never backed off and never reached MAX_QUEUE_DELIVERIES — an
unbounded flat-1s loop if step_started kept failing. It now throws, so the
error flows through the replay loop's retryable-world-error rethrow and earns
both the delivery-count backoff and the max-delivery cap. Throwing is safe on
step_started (the body hasn't run; a write that landed dedupes to skipped).
Also in this revision:
- Backoff that lasts: raise the queue handler-error retry ceiling 60s -> 900s.
VQS clamps each redelivery to its 900s SQS limit and adds its own post-32
exponential, so ramping our base toward 900s stretches survival from ~3.7h to
most of the 24h message-visibility window. Corrected the stale
MAX_QUEUE_DELIVERIES comment to match the real VQS schedule.
- Stop amplifying firewall challenges: the undici RetryAgent no longer retries
429 in-process (a challenge is a 429 the client can't solve). 429s surface
immediately as ThrottleError carrying x-vercel-mitigated / x-vercel-id, so
the diagnostic header now reaches us for the challenge case too.
- Track world faults as WORLD_CONTRACT_ERROR (not USER_ERROR) in
classifyRunError so an outage isn't attributed to user code.
- Fix queue.test.ts mock that `biome check --write` had rewritten from a
newable `function` into an arrow (broke `new QueueClient`); pin with a
biome-ignore.
All Vercel-specific logic stays in @workflow/world-vercel; @workflow/core
operates only on the generic WorkflowWorldError abstraction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Update .changeset/transport-error-redrive.md
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
* Trim changeset to a single sentence per review
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Route firewall-challenge 429s to the retryable transport path, not ThrottleError
A 429 carrying `x-vercel-mitigated: challenge` is a firewall challenge our
server-to-server client cannot solve, so it recurs for the life of the
incident. Mapping it to `ThrottleError` meant the `step_started` write deferred
it as `{ type: 'throttled' }`, which self-enqueues a FRESH queue message and
resets the delivery count — so it never backed off past `retryAfter` and never
reached `MAX_QUEUE_DELIVERIES`, hot-looping against an already-overloaded
firewall (the exact amplification this PR set out to remove, and contrary to
the "step_started can't loop unbounded" invariant, which only held for 5xx).
Map a challenge to a retryable transport `WorkflowWorldError` (`code:
'TRANSPORT'`) in both the v3 `makeRequest` and v4 `throwForErrorResponse`
(the hot event-write path) error mappings, via a shared `isFirewallChallenge429`
helper. It then propagates through the V1/V2 step paths and the replay loop's
retryable-world-error rethrow, earning the delivery-count backoff AND the
delivery cap. A genuine application-level 429 (no `challenge` mitigation) stays
a `ThrottleError` and keeps its `Retry-After`-paced defer.
Also correct the survival-window comments: with the 900s ceiling,
MAX_QUEUE_DELIVERIES=48 spans ~9-10h (~35,000s), not "the better part of 24h";
reaching 24h would need a higher delivery cap, not a higher per-hop ceiling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 41e2ad0

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 17 packages
NameType
@workflow/world-vercelPatch
@workflow/corePatch
@workflow/cliPatch
@workflow/webPatch
@workflow/buildersPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

@pranaygp
pranaygp merged commit 67fcf1a into stableJun 29, 2026
21 checks passed
@pranaygp
pranaygp deleted the backport/pr-2445-to-stable branch June 29, 2026 20:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@pranaygp