[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment - #3080

Merged
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants
Jul 24, 2026
Merged

[benchmarks/ci] SO payload variants + restructured E2E Test Results comment#3080
VaguelySerious merged 2 commits into
mainfrom
peter/bench-so-payload-variants

Conversation

@VaguelySerious

@VaguelySeriousVaguelySerious commented Jul 24, 2026

Copy link
Copy Markdown
Member

Two independent PR-comment / benchmark polish changes, both following up on #3077.

1. SO benchmark payload variants

Addresses review feedback from @karthikscale3 on #3077.

  • Deterministic, variable-length text (comment): the SO writer cycles a fixed fragment list ("The", " quick", … ".\n", ~4.5 UTF-8 bytes avg incl. punctuation/newline tokens) instead of repeating "aaaa" — closer to real token-stream traffic, still byte-for-byte reproducible.
  • Structured-delta variant (comment): a new mode: 'text' | 'structured' arg on benchSoWorkflow; structured mode wraps each fragment as { type: "text-delta", id: "0", text } (AI SDK shape). Both scenarios stream the same 300 chunks @ 100/s over 3 s and reuse the same fragments, so payload shape is the only variable and the SO gap between them isolates serialization cost.

Renders as two rows — SO / stream overhead (text) and SO / stream overhead (structured). The text scenario was renamed from stream overhead, which also drops its stale fixed-payload baseline (both SO deltas start blank and re-baseline on the next main run; no methodology bump, so other metrics keep their baselines).

2. Restructured "E2E Test Results" PR comment

Reworks .github/scripts/aggregate-e2e-results.js (aggregate mode) and .github/workflows/tests.yml:

  • Failed tests first. A renamed "Failed E2E Tests" section moves to the top and is hidden entirely when everything passed. Per category, failures under 10 are listed inline under a heading; only categories with ≥10 failures collapse into a <details>.
  • "E2E Test Summary" section follows, with two collapsibles: the overall summary table, and a flat "Details by Category" breakdown (no nested collapsibles — one big list).
  • Removed the redundant "Some E2E test jobs failed" append step (superseded by the failed-tests section) and its duplicate run link, keeping the single "View full workflow run" link from the script.

Testing

  • node --test ".github/scripts/**/*.test.js" — 29/29 pass, including new coverage for the aggregate renderer (few/many failures, all-pass, no-nested-details) and both SO rows.
  • @workflow/core typecheck passes; 97_bench.ts recompiles cleanly via SWC; both SO modes resolve through the workbench registry.
  • Aggregate comment rendering verified locally against synthetic fixtures for the <10, ≥10, and all-pass cases.
  • SO runs end-to-end in the Performance Benchmarks job on the preview deploy (the bench suite runs against Vercel only).

🤖 Generated with Claude Code

… variant
Addresses review feedback on #3077:
- The SO writer now cycles deterministic variable-length text fragments
("The", " quick", ".\n", …; ~4.5 UTF-8 bytes avg) instead of repeating
"aaaa", better approximating real token-stream traffic while staying
byte-for-byte reproducible.
- Adds a structured-delta payload mode ({ type: "text-delta", id, text })
resembling AI SDK events, run as a second SO scenario at the same 300
chunks / 100-per-sec pacing so payload shape is the only changed variable;
the SO gap between the two scenarios isolates serialization cost.
The two scenarios are labelled "stream overhead (text)" / "stream overhead
(structured)"; the rename also keeps the changed text payload from diffing
against the old fixed-payload baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8181c25

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production145502391694
✅ 💻 Local Development162102271848
✅ 📦 Local Production162102271848
✅ 🐘 Local Postgres162102271848
✅ 🪟 Windows15400154
✅ 📋 Other102002121232
✅ vercel-multi-region270027
Total7519011328651
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro126028
✅ example126028
✅ express126028
✅ fastify126028
✅ hono126028
✅ nextjs-turbopack15103
✅ nextjs-webpack15103
✅ nitro126028
✅ nuxt126028
✅ sveltekit14509
✅ vite126028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable128026
✅ express-stable128026
✅ fastify-stable128026
✅ hono-stable128026
✅ nextjs-turbopack-canary135019
✅ nextjs-turbopack-stable15400
✅ nextjs-webpack-canary135019
✅ nextjs-webpack-stable15400
✅ nitro-stable128026
✅ nuxt-stable128026
✅ sveltekit-stable14707
✅ vite-stable128026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack15400

✅ 📋 Other

AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128026
✅ e2e-local-dev-tanstack-start-128026
✅ e2e-local-postgres-nest-stable128026
✅ e2e-local-postgres-tanstack-start-128026
✅ e2e-local-prod-nest-stable128026
✅ e2e-local-prod-tanstack-start-128026
✅ e2e-vercel-prod-nest126028
✅ e2e-vercel-prod-tanstack-start126028

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@github-actions

github-actionsBot commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 8181c25 · Fri, 24 Jul 2026 00:57:48 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1262 (+32%) 🔻1351 🔴 (+26%) 🔻1427 🔴 (+26%) 🔻1814 🔴 (+27%) 🔻30
TTFSstream1275 (+32%) 🔻1349 🔴 (+29%) 🔻1377 🔴 (+29%) 🔻1509 🔴 (+33%) 🔻30
TTFShook + stream1242 (+2.3%)1683 🔴 (+29%) 🔻1756 🔴 (+29%) 🔻2626 🔴 (+86%) 🔻30
STSO1020 steps (1-20)198 (+14%)321 🔴 (+30%) 🔻364 🔴 (+30%) 🔻412 🔴 (+20%) 🔻19
STSO1020 steps (101-120)215 (+14%)288 🔴 (+4.3%)385 🔴 (+17%) 🔻436 🔴 (+16%) 🔻19
STSO1020 steps (1001-1020)492 (+7.7%)560 🔴 (+8.7%)578 🔴 (+8.6%)643 🔴 (+3.4%)19
WO1020 steps414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)414007 (+7.1%)1
SLstream latency122 (+53%) 🔻170 🔴 (+29%) 🔻192 🔴 (+7.3%)786 🔴 (+110%) 🔻30
SOstream overhead (text)13419219721730
SOstream overhead (structured)13019120032430
ℹ️ Metric definitions & methodology

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

- Move failed tests to the top as "Failed E2E Tests", hidden entirely when
everything passed. Each category lists its failures inline under a heading;
only categories with >=10 failures collapse into a <details>.
- Group the rest under an "E2E Test Summary" section with two collapsibles:
the overall summary table, and a flat "Details by Category" breakdown (no
nested collapsibles).
- Drop the redundant "Some E2E test jobs failed" append step in tests.yml
(superseded by the failed-tests section) and its duplicate workflow-run
link, keeping the single "View full workflow run" link from the script.
- Add unit coverage for the aggregate renderer (.github/scripts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@VaguelySerious
VaguelySerious marked this pull request as ready for review July 24, 2026 00:58
@VaguelySerious
VaguelySerious requested review from a team and ijjk as code ownersJuly 24, 2026 00:58
@VaguelySerious
VaguelySerious merged commit 5992507 into mainJul 24, 2026
174 of 177 checks passed
@VaguelySerious
VaguelySerious deleted the peter/bench-so-payload-variants branch July 24, 2026 02:52
@github-actions

Copy link
Copy Markdown
Contributor

Backport to stable failed for 5992507 due to a workflow error (backport job run).

This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details.

Once the underlying issue is fixed, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

599250771d1006bd59281b3346c41dfbeab22a94

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@VaguelySerious@karthikscale3