Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [workflow] Load per-source replay bundles by NathanColosimo · Pull Request #3551 · vercel/workflow · GitHub
Skip to content

[workflow] Load per-source replay bundles - #3551

Open
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles
Open

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 25 commits into
mainfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

High Level

Change build output to be separate workflow bundles and load only the needed code in the runtime.
There's very few changes to the runtime code, most of the changes are in the builder code.
Almost all of the changes are tests

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-botBot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 770e721

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/buildersPatch
@workflow/corePatch
@workflow/nextPatch
@workflow/world-testingPatch
workflowPatch
@workflow/astroPatch
@workflow/cliPatch
@workflow/nestPatch
@workflow/nitroPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated
example-nextjs-workflow-turbopackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-nextjs-workflow-webpackReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
example-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-astro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-express-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-fastify-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-hono-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nestjs-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nitro-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-nuxt-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-python-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-sveltekit-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-tanstack-start-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workbench-vite-workflowReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-docsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-swc-playgroundReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-tarballsReadyReadyPreview, v0Sep 1, 2026 3:32am UTC
workflow-webReadyReadyPreview, v0Sep 1, 2026 3:32am UTC

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3 | 🔍 observability
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M | 🔍 observability
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23 | 🔍 observability
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115 | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M1DGN5VWSB6FR072SCMVKD93
  • promiseAllWorkflow | wrun_41M1DGAP800GJNMEJ7SKXS2WR3
  • sleepingWorkflow | wrun_41M1DGBBNM0GRR3VT84RPGT37M
  • parallelSleepWorkflow | wrun_41M1DGBBX30GJQK8GDEKASRB23
  • nullByteWorkflow | wrun_41M1DGBJ7S0GGY055R4SKW2KGG
  • cancelRun - cancelling a running workflow | wrun_41M1DGG7830GNN6EDY1MQCQ115
  • cancelRun via CLI - cancelling a running workflow | wrun_41M1DGG8BD0GJG1C1TR3XYBKQP
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M1DGGFE30GK2S2QVN4AV27AY
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M1DGH1J20GQM56QKA01JAGWF

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • promiseAllWorkflow (nest)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

31 infra events
  • cold-start-warmup · suite warmup (python) · at 03:32:38Z · abandoned wrun_41M1DG9NKP0GT49DFD2FT24FQP · (+7 more)
  • run-pickup-stall · nullByteWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA4A0GWNBVG1J3PNW2BM
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F7
  • run-pickup-stall · sleepingWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA470GQ7T8KYJDV9H8F6
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDA410GG3J9DQS5J8V5RX
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:32:54Z · abandoned wrun_41M1DGDACG0GYN2PV31FE1722G
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:33:26Z · abandoned wrun_41M1DGE9JE0GYK5Q2XK9JWVR87
  • run-pickup-stall · sleepingWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5DP0GZT0NGX0185G3X8
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5BE0GMZ14CQP9SC1Y7K
  • run-pickup-stall · nullByteWorkflow (python) · at 03:33:54Z · abandoned wrun_41M1DGF5AD0GMMXJ2AJ6AZN3XB
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:33:55Z · abandoned wrun_41M1DGF5BM0GJKW6R6A11PP5G5
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:34:45Z · abandoned wrun_41M1DGF8SE0GTQDHMN40HTGC3X
  • cold-start-warmup · suite warmup (python) · at 03:35:40Z · abandoned wrun_01M1DGF7EGPY04C8PJ02V841CN · (+7 more)
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJV4MD7KSP260W7ZSBC
  • run-pickup-stall · sleepingWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK0G2QC0ZT6CH9PZYC0
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWJW8G1Y4FFC31VK77KY
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK38F43AA3AECZXYP9H
  • run-pickup-stall · nullByteWorkflow (python) · at 03:35:55Z · abandoned wrun_01M1DGJWK5160SZJ5HM2H6FMB9
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:36:26Z · abandoned wrun_41M1DGH6TB0GZCNX8E7RCMJJ6N
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:36:26Z · abandoned wrun_41M1DGJ0P00GPSJ9R5XE1S1WHX
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6HAQR493S719104TSK
  • run-pickup-stall · sleepingWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6NBAJ7B203HQDK5N2N
  • run-pickup-stall · promiseAllWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6KR8J5R749MV6Q056B
  • run-pickup-stall · parallelSleepWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ6SZKHYHZ67N6HHVCTV
  • run-pickup-stall · nullByteWorkflow (python) · at 03:36:55Z · abandoned wrun_01M1DGMQ70B63HA9BG628J3E8C
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHT45WP6FSWB7FGZRVDY
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTEREXVJRZ204H80RN5
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:37:55Z · abandoned wrun_01M1DGPHTKKDTFSEZV6GQKNE3R
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF60XCRG4QP1WMRPCM2K
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 03:38:26Z · abandoned wrun_01M1DGQF5YJJWK6M7RQPW03HKJ
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 03:38:55Z · abandoned wrun_01M1DGRCDP4DNKXYYTR30VTRPT

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production357087424320
✅ 💻 Local Development392205584480
✅ 📦 Local Production392205584480
✅ 🐘 Local Postgres392205584480
✅ 🪟 Windows32000320
❌ 🌐 Cross-language Conformance09132141
✅ vercel-http-transport8170143960
✅ vercel-multi-region270027
✅ vercel-ws-transport553087640
Total1705317277819848
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node132028
✅ astro-quickjs132028
✅ example-node132028
✅ example-quickjs132028
✅ express-node132028
✅ express-quickjs132028
✅ fastify-node132028
✅ fastify-quickjs132028
✅ hono-node132028
✅ hono-quickjs132028
✅ nest-node132028
✅ nest-quickjs132028
✅ nextjs-turbopack-node15703
✅ nextjs-turbopack-quickjs15703
✅ nextjs-webpack-node15703
✅ nextjs-webpack-quickjs15703
✅ nitro-node132028
✅ nitro-quickjs132028
✅ nuxt-node132028
✅ nuxt-quickjs132028
❌ python-node08152
✅ sveltekit-node15109
✅ sveltekit-quickjs15109
✅ tanstack-start-node132028
✅ tanstack-start-quickjs132028
✅ vite-node132028
✅ vite-quickjs132028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node134026
✅ astro-stable-quickjs134026
✅ express-stable-node134026
✅ express-stable-quickjs134026
✅ fastify-stable-node134026
✅ fastify-stable-quickjs134026
✅ hono-stable-node134026
✅ hono-stable-quickjs134026
✅ nest-stable-node134026
✅ nest-stable-quickjs134026
✅ nextjs-turbopack-canary-node141019
✅ nextjs-turbopack-canary-quickjs141019
✅ nextjs-turbopack-stable-node16000
✅ nextjs-turbopack-stable-quickjs16000
✅ nextjs-webpack-canary-node141019
✅ nextjs-webpack-canary-quickjs141019
✅ nextjs-webpack-stable-node16000
✅ nextjs-webpack-stable-quickjs16000
✅ nitro-stable-node134026
✅ nitro-stable-quickjs134026
✅ nuxt-stable-node134026
✅ nuxt-stable-quickjs134026
✅ sveltekit-stable-node15307
✅ sveltekit-stable-quickjs15307
✅ tanstack-start-node134026
✅ tanstack-start-quickjs134026
✅ vite-stable-node134026
✅ vite-stable-quickjs134026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node16000
✅ nextjs-turbopack-quickjs16000

❌ 🌐 Cross-language Conformance

AppPassedFailedSkipped
❌ python09132

✅ vercel-http-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ hono132028
✅ nextjs-turbopack15703
✅ nitro132028
✅ vite132028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

✅ vercel-ws-transport

AppPassedFailedSkipped
✅ example132028
✅ express132028
✅ nextjs-turbopack15703
✅ vite132028

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 770e721 · Tue, 01 Sep 2026 03:46:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1102 (+502%) 🔻1224 🔴 (+15%) 🔻1258 🔴 (+13%)1306 🔴 (-8.5%)30
TTFSstream112 (-48%) 💚1187 🔴 (+9.2%)1225 🔴 (+12%)1243 🔴 (+6.9%)30
TTFShook + stream377 (-69%) 💚1463 🔴 (+9.8%)1836 🔴 (+32%) 🔻2132 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)508 (-15%) 💚1679 (+5.7%)1705 (-3.1%)1752 (-6.7%)10
Fan-out TTLSPromise.all(100 steps)1590 (-4.6%)3569 (+6.8%)3692 (+4.0%)4409 (-34%) 💚10
STSO1020 steps (inline)67 (-40%) 💚133 (-10%)155 (-11%)230 (-35%) 💚1019
WO1020 steps131497 (-12%)131497 (-12%)131497 (-12%)131497 (-12%)1
CRTTfirst chunk (pooled)88 (-9.3%)120 (-20%) 💚178 (+3.5%)355 (+39%) 🔻28

Streams

ScenarioCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)110 (-25%)120 (-33%)185 (-52%)315 (-58%)116 (-40%)10
size sweep (100/s, 160B-12KB)106 (-17%)132 (-29%)268 (-21%)767 (+62%)133 (-13%)10
replay gateway-gpt-5.4-nano-2000t (1x)117 (-11%)110 (-31%)130 (-48%)224 (-63%)251 (-24%)3
replay eve-gpt-5.6-sol-2000t (1x)115 (-32%)121 (-29%)150 (-37%)335 (-61%)392 (-36%)2
replay eve-gpt-5.6-sol-2000t (2x)120 (-6%)149 (-38%)225 (-36%)457 (-28%)254 (-11%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 131259ms (Δ -17531ms, -12%)

 50-100 ms ┃ main 0 this 3 +3
100-150 ms █████████████████████░░┃ main 779 this 872 +93
150-200 ms ███┃█ main 181 this 128 -53
200-250 ms ┃ main 31 this 9 -22
250-300 ms ┃ main 11 this 6 -5
300-350 ms ┃ main 6 this 0 -6
350-400 ms ┃ main 6 this 0 -6
400-450 ms ┃ main 1 this 1 +0
450-500 ms ┃ main 1 this 0 -1
500-550 ms ┃ main 1 this 0 -1
600-650 ms ┃ main 1 this 0 -1
650-700 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ·····▁█▇▁···· 103.2 (-32%) 96 (-24%) 185 (-52%) 315 (-58%) 3000
sweep ······▆█▁▁··· 126.4 (-17%) 107 (-15%) 268 (-21%) 767 (+62%) 3000
gw 1x ·····▁█▅▁···· 95.8 (-30%) 92 (-23%) 130 (-48%) 224 (-63%) 5295
eve 1x ·····▁█▇▁▁··· 103.6 (-35%) 94 (-27%) 150 (-37%) 335 (-61%) 5186
eve 2x ·····▁▅█▂▁··· 128.1 (-32%) 113 (-31%) 225 (-36%) 457 (-28%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▂▃▂▁▅▂▁▃▃ 95–124ms
sweep █▅▃▂▂▄▂▄▂▁ 105–169ms
gw 1x ▃█▃▄▂▂▂▁▂▄ 91–107ms
eve 1x ▃▂▃▃▁█▃▄▂▁ 86–146ms
eve 2x ▂▃▃▁▁▃▆█▅▂ 103–175ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▅█▆▃▂▁▂ 123–131ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▅▃▄█▃▅▅▁▃▆ 26–35ms
sweep ▅▆▇▄▇▄▆█▄▁ 41–56ms
gw 1x ▁█▃▃▃▄▆▅▄▅ 23–32ms
eve 1x ▂▃▁▁▂█▅▅▃▃ 19–29ms
eve 2x ▂▄▁▁▁▂█▄▂▆ 19–26ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actionsBot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@vercelvercelBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional Suggestion:

Test asserts SPEC_VERSION_CURRENT === SPEC_VERSION_SUPPORTS_COMPRESSION (5) but the PR bumped SPEC_VERSION_CURRENT to SPEC_VERSION_SUPPORTS_SLOT_IDENTITY (6), so the assertion evaluates 6 === 5 and fails.

Fix on Vercel

@github-actions

github-actionsBot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
FrameworkCold replayStep reg.Framework output
hono304.7 KiB (+103.8 KiB) ⚠️40.6 KiB (±0)2.58 MiB (+830.6 KiB)
nextjs-turbopack305.0 KiB (+98.7 KiB) ⚠️439 B (±0)562.1 KiB (-203.4 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Cold replay and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.
⚠️ marks growth past the threshold. Add the allow-bundle-size-growth label to accept it.

770e721 · run

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

allow-bundle-size-growthAccept intentional bundle-size growth for this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo