Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
[world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [world] Expose capabilities in deployment health checks by NathanColosimo · Pull Request #3425 · vercel/workflow · GitHub
Skip to content

[world] Expose capabilities in deployment health checks - #3425

Open
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities
Open

[world] Expose capabilities in deployment health checks#3425
NathanColosimo wants to merge 1 commit into
mainfrom
codex/atomic-start-capabilities

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Summary

  • define WorldCapabilitiesSchema with Zod and infer WorldCapabilities from it
  • return a World's capabilities from queue-based health checks
  • keep the direct HTTP health probe lightweight and independent of World initialization
  • preserve compatibility with older and malformed health responses by treating missing or invalid capabilities as unsupported

This gives start() one capability shape for both the current World and a target deployment's World.

Stack

  1. [world] Expose capabilities in deployment health checks #3425 — World capabilities in health checks
  2. [core] Add atomic start Hook admission #3426 — Atomic start Hook admission

Testing

  • pnpm --filter @workflow/world build
  • pnpm --filter @workflow/core build
  • pnpm --filter @workflow/core exec vitest run src/runtime/helpers.test.ts
  • targeted local E2E for the direct HTTP health endpoint

@changeset-bot

changeset-botBot commented Aug 10, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 459e34b

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 20 packages
NameType
@workflow/coreMinor
@workflow/worldMinor
workflowMinor
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/world-localPatch
@workflow/world-postgresPatch
@workflow/world-vercelPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch
@workflow/nuxtPatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercelBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

ProjectDeploymentActionsUpdated (UTC)
example-nextjs-workflow-turbopackReadyReadyPreviewAug 11, 2026 9:45pm
example-nextjs-workflow-webpackReadyReadyPreviewAug 11, 2026 9:45pm
example-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-astro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-express-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-fastify-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-hono-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nestjs-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nitro-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-nuxt-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-python-workflowErrorErrorAug 11, 2026 9:45pm
workbench-sveltekit-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-tanstack-start-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workbench-vite-workflowReadyReadyPreviewAug 11, 2026 9:45pm
workflow-docsReadyReadyPreview, v0Aug 11, 2026 9:45pm
workflow-swc-playgroundReadyReadyPreviewAug 11, 2026 9:45pm
workflow-tarballsReadyReadyPreviewAug 11, 2026 9:45pm
workflow-webReadyReadyPreviewAug 11, 2026 9:45pm

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 459e34b · Tue, 11 Aug 2026 21:59:28 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1078 (+234%) 🔻1341 🔴 (+27%) 🔻1363 🔴 (+22%) 🔻1459 🔴 (+4.4%)30
TTFSstream1270 (+699%) 🔻1328 🔴 (+23%) 🔻1343 🔴 (+23%) 🔻1381 🔴 (+12%)30
TTFShook + stream496 (-63%) 💚1695 🔴 (+16%) 🔻1734 🔴 (+16%) 🔻2375 🔴 (+51%) 🔻30
STSO1020 steps (inline)122 (-3.9%)169 (-11%)192 (-11%)501 (+18%) 🔻1019
WO1020 steps179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)179029 (-2.8%)1
SLstream latency82 (-5.7%)109 🔴 (-14%)117 🔴 (-15%)147 🔴 (-70%) 💚30
SOstream overhead (text)105 (-2.8%)152 (-24%) 💚174 (-44%) 💚353 (-26%) 💚30
SOstream overhead (structured)112 (+4.7%)150 (-15%) 💚161 (-33%) 💚285 (-35%) 💚30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 183209ms → this run 177685ms (Δ -5524ms, -3%)

 100-150 ms ██████░░░░░░░┃ main 168 this 393 +225
150-200 ms ██████████████████┃█████ main 675 this 545 -130
200-250 ms █┃███ main 129 this 54 -75
250-300 ms ┃ main 23 this 9 -14
300-350 ms ┃ main 8 this 2 -6
350-400 ms ┃ main 3 this 2 -1
400-450 ms ┃ main 5 this 2 -3
450-500 ms ┃ main 3 this 1 -2
500-550 ms ┃ main 1 this 3 +2
550-600 ms ┃ main 0 this 3 +3
650-700 ms ┃ main 1 this 0 -1
800-850 ms ┃ main 1 this 0 -1
850-900 ms ┃ main 1 this 1 +0
900-950 ms ┃ main 1 this 0 -1
950-1000 ms ┃ main 0 this 1 +1
1200-1250 ms ┃ main 0 this 1 +1
4150-4200 ms ┃ main 0 this 1 +1
4250-4300 ms ┃ main 0 this 1 +1
📜 Previous results (4)

30d3802

Mon, 10 Aug 2026 23:58:42 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep1337 (+59%) 🔻1439 🔴 (+30%) 🔻1463 🔴 (+28%) 🔻1574 🔴 (+34%) 🔻30
TTFSstream1340 (+547%) 🔻1395 🔴 (+27%) 🔻1429 🔴 (+29%) 🔻1565 🔴 (+35%) 🔻30
TTFShook + stream1378 (+10.0%)1677 🔴 (+21%) 🔻1717 🔴 (+20%) 🔻1932 🔴 (+25%) 🔻30
STSO1020 steps (inline)102 (+1.0%)144 (-2.7%)172 (+1.2%)367 (+22%) 🔻1019
WO1020 steps144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)144782 (+2.1%)1
SLstream latency106 (+22%) 🔻143 🔴 (+27%) 🔻211 🔴 (+69%) 🔻395 🔴 (+178%) 🔻30
SOstream overhead (text)125 (+14%)223 (+19%) 🔻272 (+28%) 🔻1104 🔴 (+362%) 🔻30
SOstream overhead (structured)124 (+12%)261 🔴 (+61%) 🔻488 (+164%) 🔻770 (+213%) 🔻30

7084900

Mon, 10 Aug 2026 23:16:58 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep293 (-13%)1425 🔴 (+27%) 🔻1475 🔴 (+29%) 🔻1722 🔴 (+13%)30
TTFSstream224 (-78%) 💚1404 🔴 (+34%) 🔻1463 🔴 (+39%) 🔻1765 🔴 (+63%) 🔻30
TTFShook + stream358 (-72%) 💚1694 🔴 (+22%) 🔻1852 🔴 (+26%) 🔻2020 🔴 (+20%) 🔻30
STSO1020 steps (inline)85 (-11%)157 (+4.0%)186 (+3.3%)323 (-9.0%)1019
WO1020 steps156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)156230 (+5.5%)1
SLstream latency99 (+15%) 🔻195 🔴 (+50%) 🔻239 🔴 (+37%) 🔻380 🔴 (+75%) 🔻30
SOstream overhead (text)137 (+49%) 🔻211 (+13%)235 (+9.8%)900 (+259%) 🔻30
SOstream overhead (structured)143 (+39%) 🔻238 (+52%) 🔻465 (+141%) 🔻1133 🔴 (+368%) 🔻30

a6f6825

Mon, 10 Aug 2026 21:45:10 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep251 (+21%) 🔻1422 🔴 (+31%) 🔻1455 🔴 (+29%) 🔻1613 🔴 (±0%)30
TTFSstream222 (+16%) 🔻1434 🔴 (+33%) 🔻1482 🔴 (+35%) 🔻1539 🔴 (+38%) 🔻30
TTFShook + stream418 (-55%) 💚1680 🔴 (+23%) 🔻1705 🔴 (+16%) 🔻1888 🔴 (+24%) 🔻30
STSO1020 steps (inline)101 (+31%) 🔻150 (+2.0%)174 (-1.7%)268 (-46%) 💚1019
WO1020 steps145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)145990 (-4.9%)1
SLstream latency92 (+7.0%)150 🔴 (+2.0%)184 🔴 (+14%)344 🔴 (+14%)30
SOstream overhead (text)117 (-5.6%)218 (-19%) 💚248 (-24%) 💚395 (-76%) 💚30
SOstream overhead (structured)134 (+22%) 🔻295 🔴 (+5.0%)439 (-52%) 💚906 (-55%) 💚30

984728e

Mon, 10 Aug 2026 21:02:51 GMT · run logs

vercel / nextjs-turbopack

MetricScenarioBest (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstep257 (-74%) 💚1387 🔴 (+15%) 🔻1407 🔴 (+13%)1666 🔴 (+25%) 🔻30
TTFSstream230 (-11%)1472 🔴 (+30%) 🔻1515 🔴 (+27%) 🔻1946 🔴 (+40%) 🔻30
TTFShook + stream361 (-16%) 💚1648 🔴 (+9.1%)1729 🔴 (+5.7%)2028 🔴 (-4.1%)30
STSO1020 steps (inline)107 (+16%) 🔻169 (+3.0%)208 (+8.3%)397 (-7.7%)1019
WO1020 steps167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)167752 (+5.0%)1
SLstream latency108 (-3.6%)172 🔴 (-9.0%)240 🔴 (-45%) 💚500 🔴 (-14%)30
SOstream overhead (text)130 (-5.1%)327 🔴 (+37%) 🔻735 🔴 (+98%) 🔻1605 🔴 (+27%) 🔻30
SOstream overhead (structured)140 (-7.9%)313 🔴 (+20%) 🔻390 (+3.2%)673 (-2.2%)30
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
PassedFailedSkippedTotal
✅ ▲ Vercel Production346605904056
✅ 💻 Local Development381005584368
✅ 📦 Local Production381005584368
✅ 🐘 Local Postgres381005584368
✅ 🪟 Windows31200312
✅ vercel-multi-region270027
Total152350226417499
Details by Category

✅ ▲ Vercel Production

AppPassedFailedSkipped
✅ astro-node128028
✅ astro-quickjs128028
✅ example-node128028
✅ example-quickjs128028
✅ express-node128028
✅ express-quickjs128028
✅ fastify-node128028
✅ fastify-quickjs128028
✅ hono-node128028
✅ hono-quickjs128028
✅ nest-node128028
✅ nest-quickjs128028
✅ nextjs-turbopack-node15303
✅ nextjs-turbopack-quickjs15303
✅ nextjs-webpack-node15303
✅ nextjs-webpack-quickjs15303
✅ nitro-node128028
✅ nitro-quickjs128028
✅ nuxt-node128028
✅ nuxt-quickjs128028
✅ sveltekit-node14709
✅ sveltekit-quickjs14709
✅ tanstack-start-node128028
✅ tanstack-start-quickjs128028
✅ vite-node128028
✅ vite-quickjs128028

✅ 💻 Local Development

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 📦 Local Production

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🐘 Local Postgres

AppPassedFailedSkipped
✅ astro-stable-node130026
✅ astro-stable-quickjs130026
✅ express-stable-node130026
✅ express-stable-quickjs130026
✅ fastify-stable-node130026
✅ fastify-stable-quickjs130026
✅ hono-stable-node130026
✅ hono-stable-quickjs130026
✅ nest-stable-node130026
✅ nest-stable-quickjs130026
✅ nextjs-turbopack-canary-node137019
✅ nextjs-turbopack-canary-quickjs137019
✅ nextjs-turbopack-stable-node15600
✅ nextjs-turbopack-stable-quickjs15600
✅ nextjs-webpack-canary-node137019
✅ nextjs-webpack-canary-quickjs137019
✅ nextjs-webpack-stable-node15600
✅ nextjs-webpack-stable-quickjs15600
✅ nitro-stable-node130026
✅ nitro-stable-quickjs130026
✅ nuxt-stable-node130026
✅ nuxt-stable-quickjs130026
✅ sveltekit-stable-node14907
✅ sveltekit-stable-quickjs14907
✅ tanstack-start-node130026
✅ tanstack-start-quickjs130026
✅ vite-stable-node130026
✅ vite-stable-quickjs130026

✅ 🪟 Windows

AppPassedFailedSkipped
✅ nextjs-turbopack-node15600
✅ nextjs-turbopack-quickjs15600

✅ vercel-multi-region

AppPassedFailedSkipped
✅ nextjs-turbopack2700

📋 View full workflow run

@NathanColosimoNathanColosimo changed the title feat(world): expose capabilities in health checks[world] Expose capabilities in deployment health checksAug 10, 2026
Comment threadpackages/core/src/runtime/helpers.ts Outdated
@NathanColosimo
NathanColosimoforce-pushed the codex/atomic-start-capabilities branch from 681274e to 78b2b1eCompareAugust 10, 2026 22:47
Comment threadpackages/world/src/capabilities.ts Outdated

@TooTallNateTooTallNate left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 7084900. The direction is right — one capability shape for the local World and a probed target, Zod-inferred instead of hand-rolled parsing — and the suite is green as written. But the strictness semantics are a forward-compatibility trap that will fire on the very next capability addition, and I verified it experimentally.

The problem: .strict() turns "newer peer" into "unhealthy peer".

WorldCapabilitiesSchema is .strict(), and parseHealthCheckResponse does safeParse(response)null on failure. I built the branch and ran the exact scenario:

capabilities: { hookRetention: { active: true }, slotEventIds: true }
→ safeParse success: false ("Unrecognized key: slotEventIds")
→ WHOLE RESPONSE DISCARDED (healthy / specVersion / encryptionPublicKey / hookResumeInputVersion all lost)

slotEventIds isn't hypothetical — #3389 (approved, in flight) adds exactly that field, and every world it touches will declare it. The moment both PRs are in, any deployment on the newer world reads as healthy: false ("timed out") to any prober on this build. Since routine deploys create SDK version skew between live deployments, this fires in normal operation: cross-deployment start() silently degrades to the legacy wire format, loses compression, loses the sealed-args key preload, and closes the parallel-resume gate — against a perfectly healthy target. The rejects invalid World capabilities test currently encodes this failure mode as correct.

There's also a tolerance regression on the other fields: the old parser kept healthy: true and skipped a malformed optional field (specVersion: "not-a-number"); the new safeParse discards the whole response (verified). And the PR description says "treating missing or invalid capabilities as unsupported" — the code treats them as unhealthy, which is a different and much worse thing.

Requested changes:

  1. Drop .strict() from WorldCapabilitiesSchemaand the inner hookRetention object. Zod's default strip semantics give exactly the behavior the description promises: unknown capability fields are ignored (fail-closed per capability), and a newer responder parses clean.
  2. Make capabilities unable to sink the response: capabilities: WorldCapabilitiesSchema.catch(undefined).optional() (or equivalent) so even a genuinely malformed capabilities object degrades to "unsupported" while healthy/encryptionPublicKey/versions survive. Consider the same per-field .catch(undefined) for the other optionals to restore the old parser's field-level tolerance — the probe is an optimization and should never manufacture failure.
  3. Update the tests to encode tolerance: unknown capability key → known capabilities still surfaced; malformed optional field → field omitted, healthy kept. (The current rejects … tests assert the trap.)
  4. Coordinate with #3389: it adds slotEventIds to the interface this PR deletes from interfaces.ts — whichever lands second must add the field to the schema, and the merge is a semantic conflict, not a textual one, in one direction.

Non-blocking:

  • The schema's doc comments are solid but noticeably abbreviated from the interface's (the hookResumeDedup per-lookup-attestation note and the fuller deploymentAffinity rationale didn't survive). Worth carrying the operational detail over — the schema is now the only home those docs have.
  • Changeset (minor across core/world/workflow) is right for the new export.

The refactor itself is good and #3426 presumably needs it — with strip-plus-catch semantics this becomes a clean approve.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted171.0mMISMATCH1
stale-read-equal-step-countscompleted141.0mMISMATCH1
step-vs-step-forkcompleted120msMISMATCH1
step-vs-step-fork-fencedcompleted120msMISMATCH1
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mMISMATCH1
in-flight-before-decision-countedcompleted201.0mok0
in-flight-after-decisionfailed142.0mMISMATCH1
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenariooutcomeeventsvirtreplayviolations
smoke-no-stepscompleted30msok0
smoke-one-stepcompleted60msok0
hook-at-step-startedcompleted120msok0
hook-at-step-completedcompleted120msok0
hook-at-hook-createdcompleted120msok0
deadline-hook-winscompleted71.0hok0
deadline-expirescompleted71.0hok0
long-sleepcompleted1130.0dok0
hook-never-arrivesstalled30msskipped0
step-retries-twicecompleted102.0sok0
parallel-stepscompleted90msok0
hook-on-execution-statecompleted120msok0
peek-hook-before-branchcompleted120msok0
peek-hook-after-branchcompleted120msok0
peek-hook-at-registrationcompleted120msok0
race-hook-before-probecompleted120msok0
race-hook-after-probecompleted120msok0
race-duplicate-deliverycompleted130msok0
attr-hook-before-stepcompleted110msok0
attr-hook-after-stepcompleted110msok0
attr-from-step-bodycompleted130msok0
fork-hook-after-timeoutcompleted141.0mok0
fork-hook-before-timeoutcompleted141.0mok0
count-hook-after-timeoutcompleted171.0mok0
count-hook-before-timeoutcompleted201.0mok0
stale-read-step-count-forkcompleted201.0mok0
stale-read-equal-step-countscompleted141.0mok0
step-vs-step-forkcompleted120msok0
step-vs-step-fork-fencedcompleted120msok0
fence-catches-benign-directioncompleted125msok0
in-flight-before-decisioncompleted171.0mok0
in-flight-before-decision-countedcompleted171.0mok0
in-flight-after-decisioncompleted192.0mok0
stale-read-step-count-fork-fencedcompleted201.0mok0
fork-hook-winscompleted131.0mok0
fork-timeout-winscompleted131.0mok0
unclaimed-payload-under-forkcompleted171.0mok0
claimed-payload-under-forkcompleted171.0mok0
writers-independent-step-bodiescompleted120msok0
writers-scripted-tempocompleted120msok0
cancel-mid-stepcancelled70msskipped0

Full trace: world-sim-append-only.txt

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@NathanColosimo@TooTallNate@VaguelySerious