Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [debug] Validation run: combine #2113 SDK + #447 server e6722b2 by TooTallNate · Pull Request #2146 · vercel/workflow · GitHub
Skip to content

[debug] Validation run: combine #2113 SDK + #447 server e6722b2 - #2146

Closed
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528
Closed

[debug] Validation run: combine #2113 SDK + #447 server e6722b2#2146
TooTallNate wants to merge 7 commits into
peter/sdk-event-write-casfrom
debug/validate-occ-fix-20260528

Conversation

@TooTallNate

Copy link
Copy Markdown
Member

NOT FOR MERGE. Draft PR opened only to trigger the tarballs-checks workflow so we get a tarball URL to pin the repro app to.

Purpose

End-to-end validation that the combined fix closes the production-visible defect end to end. Pairs:

What this branch adds on top of #2113

  • WORKFLOW_SERVER_URL_OVERRIDE pinned to the Version Packages (beta) #447 preview URL
  • x-vercel-protection-bypass header forwarded from WORKFLOW_VERCEL_PROTECTION_BYPASS env var (so the repro app can hit the preview through Vercel Deployment Protection)

Both changes are gated to this branch only and will not be cherry-picked into either real PR.

Validation plan

Run the standard stress repro shape against the pinned tarball + preview pair (40 cycles × 200 workflows). Classify outcomes across:

  • completed
  • still running at final check
  • failed: CORRUPTED_EVENT_LOG
  • failed: USER_ERROR
  • failed: WORLD_CONTRACT_ERROR
  • failed: other

Last run pre-#447-server-fix: ~2/40 cycles surfaced CORRUPTED_EVENT_LOG on stable; 0/40 with this PR's predecessor against an earlier #447 preview but with 132 stuck-running + 23 USER_ERROR + 4 WORLD_CONTRACT_ERROR uncategorized.

Goal of this run: confirm not just CORRUPTED_EVENT_LOG = 0 but also stuck/USER_ERROR/WORLD_CONTRACT_ERROR are clean, since those would be the symptom of the materialization-before-fence orphan scenarios Peter walked.

@changeset-bot

changeset-botBot commented May 28, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7613014

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercelBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented May 28, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1190322191441
✅ 💻 Local Development161502191834
✅ 📦 Local Production161502191834
✅ 🐘 Local Postgres161502191834
✅ 🪟 Windows13100131
❌ 📋 Other7392176917
Total69053410527991

❌ Failed Tests

▲ Vercel Production (32 failed)

astro (1 failed):

  • AbortController abortExternalSignalWorkflow: signal passed as workflow input

example (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

express (3 failed):

  • outputStreamWorkflow positive startIndex (skips first chunk)
  • AbortController abortReasonWorkflow: abort reason preserved across boundaries
  • AbortController abortVoidSleepTimeoutWorkflow: documented void sleep().then(abort) pattern works

fastify (2 failed):

  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_01KSSC42D9CV29AQTYQ5VEJM1B | 🔍 observability
  • AbortController abortSurvivesReplayWorkflow: controller state consistent across replay

hono (3 failed):

  • parallelSleepWorkflow | wrun_01KSSBMQR32XYATD4XZD0KQF1A | 🔍 observability
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • spawnWorkflowFromStepWorkflow - spawning a child workflow using start() inside a step | wrun_01KSSBYRZW4NC1XCQVYX1BGMES | 🔍 observability

nextjs-turbopack (4 failed):

  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • Calculator.calculate - static workflow method using static step methods from another class | wrun_01KSSC05YKKWQY506Y50YTFD9R | 🔍 observability
  • errorSubclassRoundTripWorkflow - first-class Error subclasses survive every serialization boundary | wrun_01KSSC1XG3F71GH2WHPDG25YN8 | 🔍 observability

nextjs-webpack (1 failed):

  • AbortController abortFromStepWorkflow: step abort cancels an in-flight sibling step

nitro (1 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

nuxt (7 failed):

  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability
  • health check (queue-based) - workflow endpoint responds to health check messages
  • health check (CLI) - workflow health command reports healthy endpoints
  • pathsAliasWorkflow - TypeScript path aliases resolve correctly | wrun_01KSSBZYTZQ54CV3EQMZKKVRNT | 🔍 observability
  • AbortController abortAfterCompletionWorkflow: abort after step completes is a no-op
  • AbortController abortExternalSignalWorkflow: signal passed as workflow input
  • AbortController abortAnyInWorkflowWorkflow: AbortSignal.any composes signals inside the workflow VM

sveltekit (4 failed):

  • DurableAgent e2e core single tool call
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • closureVariableWorkflow - nested step functions with closure variables | wrun_01KSSBYPFQB35J48R1VYGS7Y9E | 🔍 observability
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC | 🔍 observability

vite (5 failed):

  • readableStreamWorkflow | wrun_01KSSBG7SCK2VJYZ4HGC1Q6H1K | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_01KSSBH63H4Y656X71066ZKSFY | 🔍 observability
  • runClassSerializationWorkflow - Run instances serialize across workflow/step boundaries | wrun_01KSSBXVA9MZPQCBQ8CTKQY50Z | 🔍 observability
  • sleepWithSequentialStepsWorkflow - sequential steps work with concurrent sleep (control) | wrun_01KSSC4D3MHRF84CKJDYMSRCZQ | 🔍 observability
  • AbortController abortTimeoutWorkflow: timeout cancels long-running step
📋 Other (2 failed)

e2e-vercel-prod-tanstack-start (2 failed):

  • stepWinsRaceWorkflow | wrun_01KSSBMZWHM1PN5WRWBSRGR530
  • fibonacciWorkflow - recursive workflow composition via start() | wrun_01KSSBZ965W97AFDHKER08Y5HC

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
❌ astro104126
❌ example104126
❌ express102326
❌ fastify103226
❌ hono102326
❌ nextjs-turbopack12542
❌ nextjs-webpack12812
❌ nitro104126
❌ nuxt98726
❌ sveltekit12047
❌ vite100526
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable106025
✅ express-stable106025
✅ fastify-stable106025
✅ hono-stable106025
✅ nextjs-turbopack-canary112019
✅ nextjs-turbopack-stable-lazy-discovery-disabled13100
✅ nextjs-turbopack-stable-lazy-discovery-enabled13100
✅ nextjs-webpack-canary112019
✅ nextjs-webpack-stable-lazy-discovery-disabled13100
✅ nextjs-webpack-stable-lazy-discovery-enabled13100
✅ nitro-stable106025
✅ nuxt-stable106025
✅ sveltekit-stable12506
✅ vite-stable106025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack13100
❌ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable106025
✅ e2e-local-dev-tanstack-start-106025
✅ e2e-local-postgres-nest-stable106025
✅ e2e-local-postgres-tanstack-start-106025
✅ e2e-local-prod-nest-stable106025
✅ e2e-local-prod-tanstack-start-106025
❌ e2e-vercel-prod-tanstack-start103226

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: success
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from c0bd797 to df40f8bCompareMay 28, 2026 22:18
Building on 98c9741's bail-on-fence-conflict, propagate the
fence-conflict signal upward as `staleSnapshot: true` so the entire
current replay's queue results are abandoned rather than just the
individual write skipped.
The narrower 'skip the write, continue the loop' shape from 98c9741
re-introduced CORRUPTED_EVENT_LOG under stress: when two concurrent
invocations make divergent VM decisions from different event-log
snapshots, the winner's fenced write succeeds and the loser's bails.
But the loser's VM had already derived its own queue results from the
stale snapshot — if it continues past the conflict and queues them,
those queue items can drive subsequent ticks that consume the
winner's events as their own, surfacing as the original step_mismatch
shape ("step_started for step_X belongs to <name-A> but consumer is
<name-B>").
The right behavior is the one Pranay sketched on Slack: 'if new events
have been introduced to the log after a concurrent replay has started,
the invocation queue results must be abandoned. That replay is invalid.'
Implementation:
- `FencedWriteResult` now carries a `staleSnapshot` boolean so callers
can distinguish 'fence conflict — abandon entire replay' from
'entity already exists — skip this write but keep going'.
- `handleSuspension` short-circuits and returns `{ staleSnapshot: true,
pendingSteps: [] }` the moment any fenced write rejects with a
fence conflict. Subsequent step/wait writes from that replay never
run.
- Runtime tick detects `staleSnapshot: true` and `return`s cleanly
(no `run_failed` event). The canonical invocation is left to make
progress; the run stays `running`.
The elapsed-wait scan (`wait_completed`) deliberately keeps its
continue-on-conflict shape: the work it derives is purely
timer-based (which waits have elapsed), not a VM branch decision, so
a stale snapshot doesn't change the set of waits to complete. Only
the suspension handler's writes are guarded by the abandon-the-tick
semantic.
Tests: 1018 core tests pass.
@TooTallNate
TooTallNateforce-pushed the debug/validate-occ-fix-20260528 branch from 455bda9 to 0c7eb75CompareMay 28, 2026 23:26
The abandon-tick change (fbaa2bf) correctly stops a stale-snapshot
replay from queueing divergent work, but it returned without
re-enqueueing. Under a hook burst, every tick that would consume the
late-arriving hook_received events could race and abandon, leaving the
run 'running' with pending hooks and no tick scheduled to advance it.
Stress testing showed ~28/40 runs stalled this way (valid fence, real
events, just no continuation).
Return { timeoutSeconds: 0 } on stale-snapshot abandon instead of a
bare return — the same immediate re-enqueue idiom the hook-conflict
path already uses. This guarantees a fresh tick re-runs against the
canonical event log.
This is bounded (one re-enqueue per abandoned tick) and converges:
paired with the server-side atomic fence+event write (no phantom
fences), the canonical replay makes forward progress, so the
re-enqueued tick advances the log rather than spinning — unlike the
original MAX_FENCE_RETRIES storm this design replaced.
The orphaned-step-dispatch recovery (re-queue step_created /
step_retrying events that never reached step_started) was gated on
`metadata.attempt > 1`, i.e. only on queue redeliveries. That misses
the stale-snapshot abandon path: when a tick writes a fenced
step_created and then abandons on a *later* fenced write (returning
staleSnapshot + re-enqueuing), it never reaches the step-queueing
code. The re-enqueue produces a *fresh* queue message (attempt 1), not
a redelivery, so the attempt-gated recovery never fired — leaving the
run stalled with a valid fence and an orphaned step_created that no
one dispatches.
Run the recovery scan on every invocation. It is safe unconditionally:
step dispatch is queued with `idempotencyKey: step.correlationId`, so
re-queueing an already-dispatched step is deduped by the queue. Steps
this tick created are still queued via `createdStepCorrelationIds` and
selected for inline execution via `ownedPendingSteps` (unchanged);
recovery only adds orphans this tick did not create, which are queued
(never inline-executed) — correct, since their creating tick abandoned.
Observed in stress testing: with the atomic-fence server fix
eliminating phantom fences, a residual set of runs stalled with a real
fence + a step_created that never started. This closes that gap.
Tests: 1018 core tests pass.
…shot replay
The previous attempt (unconditional orphaned-step recovery scan,
reverted in d441126) re-queued every pending step_created on every
invocation. That violated the single-owner-per-step invariant: a
non-owner tick could re-dispatch a step another tick was already
running, producing a duplicate step_started and a
CORRUPTED_EVENT_LOG ("Unconsumed event in event log:
eventType=step_started"). 2/40 runs hit this in stress.
Safer approach: when handleSuspension abandons on a stale snapshot, it
returns the steps it ALREADY wrote a fenced step_created for (the ones
in createdStepCorrelationIds) as pendingSteps. Those writes succeeded
against a matching fence inside the atomic transaction, so they're
canonical and owned by exactly this tick. The runtime's
staleSnapshot branch dispatches just those owned steps (with
idempotencyKey: correlationId) before re-enqueuing, so:
- no orphaned step_created (the step that this tick created always gets
an owner to dispatch it), and
- no double-dispatch (only the single owning tick queues each step;
other ticks that abandon before writing the step_created never claim
ownership of it).
This pairs ownership with dispatch instead of blindly recovering, which
is what made the unconditional scan unsafe.
Tests: 1018 core tests pass.
NOT FOR MERGE. This branch ('debug/validate-occ-fix-20260528') exists
to run end-to-end stress validation of the combined fix:
- @workflow/core: #2113 (e5cd686, top of branch)
- workflow-server: vercel/workflow-server#447 (e6722b2)
WORKFLOW_SERVER_URL_OVERRIDE is pinned to the workflow-server preview
deployment for #447 so the repro app exercises both PRs together. The
WORKFLOW_VERCEL_PROTECTION_BYPASS env var is forwarded as the bare
'x-vercel-protection-bypass' header to bypass the preview's Vercel
Deployment Protection. Setting 'x-vercel-set-bypass-cookie: true' is
deliberately NOT done — it triggers a 307 redirect loop on Node undici.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@TooTallNate