[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Closed
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Closed

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

@pranaygppranaygp commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completedrow (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_createdrow it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR (#782) heals fresh wait wedges by backfill and cancels stale ones; this PR is the SDK-side detection so the contradiction is loud while it persists — and, where a safe anchor exists, terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): warn, then fail. When the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt (durable, adopted from the wait_created row, identical on every wake), fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): warn only. No per-wait replay-stable time anchor exists for an uncreated wait, so this site never fails the run:

  • Its resumeAt is recomputed from the live clock on every replay, so it always sits in the future.
  • The ULID inside its correlation id encodes the run's creation epoch, not the wait's scheduling instant — the workflow VM mints every correlation id as ulid(fixedTimestamp) (a run-wide constant, held fixed precisely so ids are replay-stable). An earlier revision of this PR escalated on that ULID; as review pointed out, that would fail any sufficiently old run with CORRUPTED_EVENT_LOG on a single benign concurrent-suspension race.

The run epoch still works as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is untouched), and a fresh event-log read gates the warning so a concurrent winner's late-landing row is not reported as a wedge. Terminating a genuinely stale wait_created wedge is owned by workflow-server #782's stale-cancel tier, which classifies against entity state the SDK cannot see.

Why stateless, time-based escalation (where it applies): every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. For wedged completions it is the warn→fail boundary; for wedged creations it is the detection pre-filter. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The throw happens in the elapsed-wait pass of the replay loop, which already routes it to the terminal path (same as the slot-gap check) — the suspension handler no longer throws, so no error-routing changes remain in this PR.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, run-epoch ULID decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; wedged completions warn inside the threshold and fail with CORRUPTED_EVENT_LOG past it; wedged creations warn (span attribute + log) but never fail and keep the run's normal suspension behavior.
  • Full core suite: 97 files, 2134 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.
Make the contradiction loud, and terminal past a generous threshold:
- wait_completed (elapsed-wait pass): when the create conflicts AND the
follow-up reload still cannot produce the row, warn and report
workflow.wait.wedge_suspected on the invocation span; once the clock
is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
(an uncreated wait recomputes it from the live clock every replay), so
the anchor is the scheduling instant embedded in the wait's
replay-stable correlation id. Past the threshold the contradiction is
verified against a fresh event-log read before failing, so a
concurrent writer's row landing after this replay's snapshot is never
mistaken for a wedge.
Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code ownersAugust 14, 2026 01:08
CopilotAI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
workflowPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 14, 2026 1:11am
example-nextjs-workflow-webpackReadyReadyPreviewAug 14, 2026 1:11am
example-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-astro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-express-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-fastify-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-hono-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nestjs-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nitro-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-nuxt-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-python-workflowErrorErrorAug 14, 2026 1:11am
workbench-sveltekit-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-tanstack-start-workflowReadyReadyPreviewAug 14, 2026 1:11am
workbench-vite-workflowReadyReadyPreviewAug 14, 2026 1:11am
workflow-docsReadyReadyPreview, v0Aug 14, 2026 1:11am
workflow-swc-playgroundReadyReadyPreviewAug 14, 2026 1:11am
workflow-tarballsReadyReadyPreviewAug 14, 2026 1:11am
workflow-webReadyReadyPreviewAug 14, 2026 1:11am

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development353605204056
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total149610222617187
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1076 (+182%) 🔻1347 🔴 (+21%) 🔻1375 🔴 (+21%) 🔻1405 🔴 (-8.3%)30
TTFSstream1285 (+28%) 🔻1340 🔴 (+27%) 🔻1351 🔴 (+26%) 🔻1471 🔴 (+33%) 🔻30
TTFShook + stream1553 (+22%) 🔻1635 🔴 (+18%) 🔻1650 🔴 (+16%) 🔻1703 🔴 (+5.1%)30
Fan-out TTFSPromise.all(100 steps)8809 (-1.2%)10727 (+7.8%)14678 (+46%) 🔻15082 (+12%)10
Fan-out TTLSPromise.all(100 steps)17196 (-2.7%)20436 (+8.3%)23214 (+22%) 🔻23631 (+0.8%)10
STSO1020 steps (inline)124 (+0.8%)174 (-8.9%)197 (-14%)391 (-33%) 💚1019
WO1020 steps171617 (-12%)171617 (-12%)171617 (-12%)171617 (-12%)1
SLstream latency78 (-1.3%)105 🔴 (-4.5%)117 🔴 (-9.3%)129 🔴 (-62%) 💚30
SOstream overhead (text)97 (-13%)141 (-22%) 💚157 (-24%) 💚314 (-48%) 💚30
SOstream overhead (structured)101 (+5.2%)154 (-1.3%)194 (+16%) 🔻8083 🔴 (+4341%) 🔻30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

 100-150 ms ███████░░░┃ main 180 this 295 +115
150-200 ms ███████████████████████┃ main 627 this 631 +4
200-250 ms █┃███ main 134 this 57 -77
250-300 ms ┃ main 29 this 14 -15
300-350 ms ┃ main 15 this 7 -8
350-400 ms ┃ main 11 this 5 -6
400-450 ms ┃ main 4 this 5 +1
450-500 ms ┃ main 5 this 3 -2
500-550 ms ┃ main 3 this 0 -3
550-600 ms ┃ main 1 this 2 +1
600-650 ms ┃ main 5 this 0 -5
650-700 ms ┃ main 1 this 0 -1
750-800 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
1100-1150 ms ┃ main 1 this 0 -1
4450-4500 ms ┃ main 1 this 0 -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
packages/core/src/telemetry/semantic-conventions.tsAdds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.tsNew wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.tsUnit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.tsEnd-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.tsAdds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.tsAdds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdxDocuments WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.mdChangeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Restructured in ee40e67. For the record, the aliased-condition form did typecheck (TS 4.4+ narrows const booleans built from narrowing conjunctions), but the fragility point was fair — and the rework for the other review thread rebuilt this branch anyway: the !== undefined check is now part of the if condition directly, so the narrowing is structural rather than aliased, and the timestamp is no longer used in the body beyond the guard.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisionfailed91.0mMISMATCH1
in-flight-before-decision-countedfailed91.0mMISMATCH1
in-flight-after-decisionfailed91.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

@vercelvercelBotAug 14, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in ee40e67 — this finding was correct. generateUlid is ulid(fixedTimestamp) (workflow.ts), where fixedTimestamp is the run's creation epoch held constant for replay stability, so the correlation-id ULID is a run-wide constant and cannot anchor a per-wait escalation: any run older than the threshold would fail as CORRUPTED_EVENT_LOG on one benign concurrent-suspension race whenever the winner's row wasn't yet visible to the verification read.

The wait_created site is now warn-only: the run-epoch timestamp is kept solely as a cheap pre-filter (conflicts in runs younger than the threshold skip detection entirely, so the hot path is unchanged), the fresh-log-read verification still gates the warning, and escalation for genuinely stale wait_created wedges is left to workflow-server#782's recovery (backfill fresh, cancel stale), which classifies against entity state the SDK cannot see. The wait_completed site keeps warn→fail — its resumeAt anchor is durable and genuinely per-wait. Docs, tests, and the PR description are updated to match.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming and fixing. Removing the ULID-anchored escalation for the wait_created site is the right call — since generateUlid is ulid(fixedTimestamp) with a run-wide constant, there's no sound per-wait anchor available there, and the correctly-anchored wait_completed path (on the durable resumeAt) is unaffected. The tradeoff — a genuine wait_created wedge reverts to the pre-PR silent wake-loop rather than being detected — is strictly safer than failing healthy long-running runs; sound detection would need a durable per-wait anchor (e.g. a World read of the wait entity).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@pranaygp@VaguelySerious