[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main - #3213

Merged
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1
Jul 30, 2026
Merged

[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main#3213
shalabhc merged 5 commits into
mainfrom
shalabhc/benchmark-variance-fix1

Conversation

@shalabhc

@shalabhcshalabhc commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

Why

The sequential-steps benchmark's STSO metric mixed two unrelated phenomena:

  • inline gaps — steps running back-to-back in the same warm process (pure framework overhead, ~200-500ms)
  • queue-hop gaps — steps across an invocation boundary, paying queue dispatch + client reinit + event-log replay (~2-3s, roughly 10x more)

The old step-index windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured neither cleanly. A 1020-step run crosses ~3 invocation boundaries at non-deterministic step indices, so whether a boundary happened to land inside a window moved that window's P99 by hundreds of percent between two runs of the same commit against the same deployment. That was most of the run-to-run variance the benchmark comment has been reporting.

Diagnosed on shalabhc/benchmarks-variance, where separating the two categories showed inline STSO to be stable within ~3% across runs while the visible "regressions" were entirely boundary placement.

What changed

Ground-truth tagging.97_bench.ts tags each step with whether it was the first step body executed in its process (queue-hop) or a later one in the same warm process (inline), via a process-global. Not inferred from step index or trace timestamps.

Two STSO rows over every gap instead of three sampled 19-gap windows. No targets on the new rows — the old ones described the index-bucketed grouping.

Full sample retention.computeStats keeps the sorted sample array alongside the percentiles.

Histogram + cumulative-time diff vs main, one per STSO kind, rendered under the existing table (which is unchanged). Percentiles hide how many samples moved and by how much, which is exactly where the variance lives.

No changes to benchmarks.yml — this is still one run diffed against the main baseline the workflow already downloads.

Sample output

**STSO distribution vs `main`**
_1020 steps (inline)_
Cumulative STSO time: main 369664ms → this run 333174ms (Δ -36490ms, -10%)
200-250 ms █┃ main 2 this 35 +33
250-300 ms ███░░░░░░░░┃ main 59 this 227 +168
300-350 ms █████████████████░░░░░░┃ main 322 this 445 +123
350-400 ms █████████████┃█████████ main 433 this 254 -179
400-450 ms ██┃███████ main 185 this 54 -131
450-500 ms ┃ main 15 this 1 -14

is main, marks where this run lands, bridges the gap when this run has more samples in a bucket. Per-bucket counts and deltas ride on the bar line — no second table restating them.

Notes for review

  • This PR's own comment will show the distributions without deltas. No run on main has recorded raw samples yet, so the section renders as a single series (with a note) rather than diffing this run against itself. Deltas start on the next PR.
  • Raw samples are stripped from the comment's embedded data block. ~1000 samples × current + baseline would exceed GitHub's 65k comment limit within a couple of history entries. The histogram therefore renders for the current run only; collapsed history keeps its tables. Simulated 12 successive runs → 15.9 KB comment with a full 10-entry history.
  • Inline rows use a fixed 50ms bin width; the adaptive width is coarse enough to hide structure inside that cluster (a bimodal split was invisible behind ~500ms bins). Queue-hop rows keep the adaptive width.
  • Negative gaps get their own <0 (skew) bucket. Step timestamps come from two different step bodies, so a gap can come out slightly negative under clock skew; counting those with the slowest samples inverts what the tail bucket means.
  • Empty changeset, matching [benchmarks/ci] SO payload variants + restructured E2E Test Results comment #3080 which touched the same file set.

Draft until the benchmark job runs on it and the rendered comment is confirmed.

🤖 Generated with Claude Code

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@changeset-bot

changeset-botBot commented Jul 30, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 342f5a3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewJul 30, 2026 7:38pm
example-nextjs-workflow-webpackReadyReadyPreviewJul 30, 2026 7:38pm
example-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-astro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-express-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-fastify-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-hono-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nestjs-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nitro-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-nuxt-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-sveltekit-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-tanstack-start-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workbench-vite-workflowReadyReadyPreviewJul 30, 2026 7:38pm
workflow-docsReadyReadyPreview, v0Jul 30, 2026 7:38pm
workflow-swc-playgroundReadyReadyPreviewJul 30, 2026 7:38pm
workflow-tarballsReadyReadyPreviewJul 30, 2026 7:38pm
workflow-webReadyReadyPreviewJul 30, 2026 7:38pm

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

💻 Local Development (1 failed)

nextjs-webpack-stable (1 failed):

  • AbortController abortAnyInStepWorkflow: AbortSignal.any inside a step composes deserialized signals

📦 Local Production (1 failed)

nextjs-webpack-stable (1 failed):

  • webhookWorkflow | wrun_41KYT8GD2H0GHXMVEV7SMN9E74

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
❌ 💻 Local Development162012271848
❌ 📦 Local Production162012271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7517211328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

❌ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

❌ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
❌ nextjs-webpack-stable15310
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 342f5a3 · Thu, 30 Jul 2026 19:56:34 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep273 (-10%)1402 🔴 (+30%) 🔻1439 🔴 (+31%) 🔻1526 🔴 (+12%)30
TTFSstream258 (-74%) 💚1398 🔴 (+36%) 🔻1418 🔴 (+37%) 🔻1491 🔴 (+30%) 🔻30
TTFShook + stream345 (-72%) 💚1601 🔴 (+18%) 🔻1657 🔴 (+19%) 🔻1748 🔴 (+19%) 🔻30
STSO1020 steps (inline)1664965547361016
STSO1020 steps (queue-hop)15452632263226323
WO1020 steps423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)423331 (+6.1%)1
SLstream latency108 (+24%) 🔻185 🔴 (+22%) 🔻217 🔴 (+10%)233 🔴 (-39%) 💚30
SOstream overhead (text)142 (+22%) 🔻237 (+22%) 🔻267 (-45%) 💚369 (-51%) 💚30
SOstream overhead (structured)125 (+21%) 🔻250 (+18%) 🔻312 (+23%) 🔻480 (+6.4%)30
📈 STSO distribution (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: 416691ms over 1016 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

 150-200 ms ████ steps 26
200-250 ms ███████████ steps 67
250-300 ms ███████████████████ steps 117
300-350 ms ████████████████████ steps 123
350-400 ms ████████████████████████ steps 148
400-450 ms ████████████████████████ steps 150
450-500 ms ██████████████████████ steps 139
500-550 ms █████████████████████ steps 132
550-600 ms ███████████ steps 69
600-650 ms ████ steps 22
650-700 ms █ steps 7
700-750 ms █ steps 7
750-800 ms █ steps 4
800-850 ms █ steps 2
900-950 ms █ steps 1
1000-1050 ms █ steps 1
1050-1100 ms █ steps 1

1020 steps (queue-hop)

Cumulative STSO time: 6302ms over 3 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

1500-2000 ms ████████████████████████ steps 1
2000-2500 ms ████████████████████████ steps 1
2500-3000 ms ████████████████████████ steps 1
📜 Previous results (2)

1bfc037

Thu, 30 Jul 2026 19:20:17 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1275 (+423%) 🔻1336 🔴 (+22%) 🔻1371 🔴 (+21%) 🔻1408 🔴 (-18%) 💚30
TTFSstream1271 (+265%) 🔻1319 🔴 (+23%) 🔻1323 🔴 (+22%) 🔻1355 🔴 (+20%) 🔻30
TTFShook + stream1509 (+20%) 🔻1590 🔴 (+15%)1620 🔴 (+13%)1712 🔴 (+14%)30
STSO1020 steps (inline)1654665177311016
STSO1020 steps (queue-hop)22563125312531253
WO1020 steps396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚396520 (-17%) 💚1
SLstream latency96 (-1.0%)148 🔴 (+3.5%)155 🔴 (+6.2%)236 🔴 (-34%) 💚30
SOstream overhead (text)97 (-24%) 💚146 (-49%) 💚157 (-57%) 💚209 (-64%) 💚30
SOstream overhead (structured)106 (-16%) 💚159 (-50%) 💚186 (-50%) 💚214 (-62%) 💚30

95cf46a

Thu, 30 Jul 2026 18:51:40 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1267 (+473%) 🔻1327 🔴 (+25%) 🔻1346 🔴 (+22%) 🔻1431 🔴 (-6.8%)30
TTFSstream254 (+12%)1297 🔴 (+34%) 🔻1319 🔴 (+28%) 🔻1358 🔴 (+24%) 🔻30
TTFShook + stream475 (+13%)1576 🔴 (+30%) 🔻1598 🔴 (+28%) 🔻1877 🔴 (+31%) 🔻30
STSO1020 steps (inline)1704565157431016
STSO1020 steps (queue-hop)22703446344634463
WO1020 steps395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)395976 (-6.2%)1
SLstream latency91 (+7.1%)139 🔴 (+6.9%)150 🔴 (-9.6%)281 🔴 (-22%) 💚30
SOstream overhead (text)103 (-14%)151 (-34%) 💚183 (-30%) 💚231 (-35%) 💚30
SOstream overhead (structured)114 (-8.8%)162 (-35%) 💚276 (-13%)848 (+101%) 🔻30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

@karthikscale3karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark instrumentation and reporting changes. The current E2E failures are in unrelated webpack webhook/HMR tests and do not appear caused by this PR.

@VaguelySeriousVaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, but can we hide "STSO distribution" in a collapsible entry in the PR comment?

vercelBotand others added 2 commits July 30, 2026 19:34
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 8bda7ce (AI decision).

This is an enhancement to internal benchmarking tooling: it re-partitions the STSO metric into new inline/queue-hop categories, adds raw sample retention, and adds a new histogram/cumulative-time diff section to the benchmark PR comment. It adds new measurement capability and reporting surface rather than fixing a defect that affects stable, and reducing benchmark run-to-run variance is a measurement-quality improvement, not a stability fix for the maintenance line.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

8bda7cef79563a1e094e77a52dd87743db513dad

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shalabhc@VaguelySerious@karthikscale3