Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [e2e] Host the abort-fetch slow endpoint inside the step by alangenfeld · Pull Request #3618 · vercel/workflow · GitHub
Skip to content

[e2e] Host the abort-fetch slow endpoint inside the step - #3618

Merged
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch
Aug 18, 2026
Merged

[e2e] Host the abort-fetch slow endpoint inside the step#3618
alangenfeld merged 1 commit into
mainfrom
alangenfeld/e2e-hermetic-abort-fetch

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

The abort-fetch e2e tests raced a 2s sleep against a fetch to public slow endpoints (postman-echo, httpbin /delay/10), so upstream 5xxs and early returns from CI runners showed up as flakes. fetchWithSignal now stands up a loopback node:http server that holds each response open for ~30s, keeping the subject — cancelling a real in-flight fetch — with no external dependency. The 30s hold is longer than the tests' 60s budgets minus the race, so broken abort propagation still fails as a natural completion rather than a hang.

Test Plan

Existing abort-fetch e2e coverage runs in CI; workbench/example builds clean.

The abort-fetch tests cancelled an in-flight fetch against external slow
endpoints (postman-echo, httpbin /delay/10, tried in order). Those
upstreams 5xx and return early from GH Actions runners often enough to
be a recurring flake class - the tests were measuring the public
internet instead of abort propagation - and heavier suite load (e.g.
re-enabling e2e concurrency, #2083) makes both upstreams flake at once.
fetchWithSignal now hosts its own slow endpoint: an in-process node:http
server on a loopback ephemeral port that holds each response open for
~30s. The subject is unchanged - a real in-flight HTTP fetch cancelled
mid-flight - with no external dependency. The 30s hold keeps regression
detection honest: broken abort propagation surfaces as natural
completion (ok: true) within the tests' 60s budgets.
A per-workbench /api/delay route was rejected earlier because it would
only exist on whichever workbench it was added to; the in-step server
travels with the workflow fixture to every app.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@vercel

vercelBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 18, 2026 7:28pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 18, 2026 7:28pm
example-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-astro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-express-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-fastify-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-hono-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nestjs-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nitro-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-nuxt-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-python-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-sveltekit-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workbench-vite-workflowReadyReadyPreviewAug 18, 2026 7:28pm
workflow-docsReadyReadyPreview, v0Aug 18, 2026 7:28pm
workflow-swc-playgroundReadyReadyPreviewAug 18, 2026 7:28pm
workflow-tarballsReadyReadyPreviewAug 18, 2026 7:28pm
workflow-webReadyReadyPreviewAug 18, 2026 7:28pm

@changeset-bot

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d7b0736

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (1 failed)

nextjs-turbopack-quickjs (1 failed):

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • basic text response (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 19:31:37Z · abandoned wrun_01M0B5PQWJHY8A8XDHNE27EWY3

E2E Test Summary

Summary
PassedFailedSkippedTotal
❌ ▲ Vercel Production347317384212
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ 🌐 Cross-language Conformance90128137
✅ vercel-multi-region270027
Total152511254017792
Details by Category

❌ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
❌ nextjs-turbopack-quickjs15213
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ python-node80148
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ 🌐 Cross-language Conformance

AppPassedFailedSkipped
✅ python90128

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d7b0736 · Tue, 18 Aug 2026 19:46:55 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1376 (+464%) 🔻1504 🔴 (+32%) 🔻1540 🔴 (+26%) 🔻1585 🔴 (-8.8%)30
TTFSstream354 (+52%) 🔻1424 🔴 (+23%) 🔻1463 🔴 (+23%) 🔻1495 🔴 (+14%)30
TTFShook + stream1346 (+214%) 🔻1735 🔴 (+25%) 🔻1813 🔴 (+29%) 🔻1996 🔴 (+36%) 🔻30
Fan-out TTFSPromise.all(100 steps)1188 (-17%) 💚3255 (+29%) 🔻3428 (+30%) 🔻5328 (+61%) 🔻10
Fan-out TTLSPromise.all(100 steps)5371 (-25%) 💚7805 (-11%)10551 (+17%) 🔻10957 (+21%) 🔻10
STSO1020 steps (inline)122 (-25%) 💚167 (-44%) 💚204 (-42%) 💚412 (-25%) 💚1019
WO1020 steps170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚170113 (-41%) 💚1
CRTTfirst chunk (pooled)113 (+14%)166 (-1.8%)375 (+15%)3821 (+738%) 🔻28

Streams

Scenariowr c/srd c/swr KiB/srd KiB/sCRTT 1stp75p90p99CDV maxiters
paced control (100/s, 60B)100 (±0%)101 (±0%)5 (±0%)5.1 (±0%)134 (-13%)172 (-21%)593 (+79%)3796 (+271%)157 (-51%)10
size sweep (100/s, 160B-12KB)100 (±0%)99.4 (-2%)334 (±0%)331 (-2%)123 (-2%)182 (-7%)234 (-76%)379 (-69%)159 (-23%)10
replay gateway-gpt-5.4-nano-2000t (1x)89.2 (±0%)89.4 (±0%)16.2 (±0%)16.3 (+1%)166 (+8%)142 (-27%)186 (-33%)402 (-50%)209 (-44%)3
replay eve-gpt-5.6-sol-2000t (1x)54.7 (±0%)54.7 (±0%)355 (±0%)355 (±0%)136 (-12%)140 (-11%)177 (-10%)326 (-21%)221 (-33%)2
replay eve-gpt-5.6-sol-2000t (2x)109 (±0%)110 (±0%)710 (±0%)712 (±0%)134 (-35%)212 (-15%)308 (-7%)1219 (+6%)305 (-54%)3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 287804ms → this run 168740ms (Δ -119064ms, -41%)

 100-150 ms ░░░░░░░░░░░░░░░░░░░░░░░┃ main 0 this 476 +476
150-200 ms █░░░░░░░░░░░░░░░░░░░░┃ main 27 this 432 +405
200-250 ms ██┃████████████ main 305 this 57 -248
250-300 ms ┃█████████████████████ main 431 this 25 -406
300-350 ms ┃██████ main 142 this 12 -130
350-400 ms ┃███ main 71 this 5 -66
400-450 ms ┃ main 11 this 4 -7
450-500 ms ┃ main 11 this 4 -7
500-550 ms ┃ main 10 this 0 -10
550-600 ms ┃ main 4 this 1 -3
600-650 ms ┃ main 1 this 1 +0
650-700 ms ┃ main 2 this 1 -1
700-750 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
1100-1150 ms ┃ main 1 this 0 -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant RTT 1ms→5s+ avg p50 p90 p99 n
control ······▂█▂▁▁▁· 372.7 (+96%) 128 (-17%) 593 (+79%) 3796 (+271%) 3000
sweep ······▃█▂···· 139.9 (-29%) 128 (-10%) 234 (-76%) 379 (-69%) 3000
gw 1x ·····▁▅█▁···· 122 (-20%) 117 (-13%) 186 (-33%) 402 (-50%) 5295
eve 1x ·····▁▆█▁···· 118.1 (-10%) 104 (-11%) 177 (-10%) 326 (-21%) 5186
eve 2x ·····▁▂█▃▁▁·· 178.7 (-13%) 144 (-17%) 308 (-7%) 1219 (+6%) 7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control █▇▆▅▄▄▂▄▃▁ 248–507ms
sweep █▇▃▅▅▅▃▂▄▁ 128–153ms
gw 1x █▃▁▄▃▂▃▂▃▃ 105–159ms
eve 1x ▃▂▄▃▁▂█▆▆▁ 105–141ms
eve 2x ▃█▂▂▁▂▄▅▂▁ 126–307ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep ▂▁▄█▆▃▁ 139–142ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control ▁▃▄▂▃▂▂█▁▂ 34–62ms
sweep ▁▇█▃▆▇▂▇▄▂ 46–61ms
gw 1x █▂▁▁▅▆▆▂▃▂ 31–42ms
eve 1x ▂▅▄▄▁▅▆▅▄█ 21–26ms
eve 2x ▅█▂▃▂▄▃▁▂▃ 20–37ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000teaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 18, 2026 19:32
@alangenfeld
alangenfeld requested a review from a team as a code ownerAugust 18, 2026 19:32

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Smart

@alangenfeld
alangenfeld merged commit 8789f45 into mainAug 18, 2026
428 of 434 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-hermetic-abort-fetch branch August 18, 2026 21:59
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8789f45 (AI decision).

This is a test-only flake fix, which would normally qualify, but the behavior it covers does not exist on stable: git show origin/stable shows neither fetchWithSignal/SLOW_FETCH_URLS nor any AbortController/abort workflows in workbench/example/workflows/99_e2e.ts, and no abort-fetch cases in packages/core/e2e/e2e.test.ts (both files exist on stable, just without this coverage). With no abort-fetch tests on the maintenance line, there is no flake there to fix.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8789f4529b311205e41f38a7f6fb188677e6dcc7

alangenfeld added a commit that referenced this pull request Aug 18, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
alangenfeld added a commit that referenced this pull request Aug 19, 2026
Serial execution has been the dominant wall-clock cost per matrix entry
since concurrency was disabled before conf (78048e0): ~128 tests at
~22 of 24 minutes on the Vercel lanes, and lately the slowest lane
cannot finish under its 30-minute job timeout on a slow runner day at
all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on
the vitest phase) and identified what broke; its blockers are now fixed:
world-local writeExclusive is atomic (write-then-link), abort-fetch
tests are hermetic (#3618), the fibonacci tree fits the scheduler
(#3619), and source-map assertions are positive-only (#3620).
What this change adds is concurrency-safe per-test attribution. The
harness tracked runs and test names in module globals reset by a
beforeEach - under concurrency every test clobbered every other's
state, so a failing test dumped an unrelated sibling's diagnostics.
vitest's getCurrentTest() cannot substitute: it is a plain module
variable, wrong after any await. Instead an auto fixture - the one
place that receives the test's own context unambiguously - binds a
per-test state (name, tracked runs, the test's own skip) via
AsyncLocalStorage around each test body, and trackRun /
recordInfraEvent / requireFixture read it ambiently with no call-site
changes. The conformance gates skip through the bound state's skip, so
a mid-body requireFixture skips the right test. Sequential suites
(dev, agent, region) keep setupRunTracking's module-global fallback.
Full suite passes 137/137 concurrently against a local dev server in
under 2 minutes. A test that genuinely cannot share a deployment can
opt out with test.sequential.
Builds on VaguelySerious's investigation in #2083.
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alangenfeld@VaguelySerious