perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

perf(core): retain workflow VM across inline steps - #2984

Closed
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm
Closed

perf(core): retain workflow VM across inline steps#2984
NathanColosimo wants to merge 4 commits into
mainfrom
codex/retain-workflow-vm

Conversation

@NathanColosimo

@NathanColosimoNathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

Summary

  • retain one workflow VM, event consumer, async stack, and hydrated state across inline step suspensions within a single queue invocation
  • append only newly durable events when the inline loop resumes the workflow
  • fall back permanently to ordinary replay if the durable event prefix changes or a discarded session advances
  • preserve runWorkflow as the one-shot compatibility API
  • tag workflow.run spans with workflow.execution.mode=replay|retained for direct production decomposition

Design

runWorkflowSession owns the invocation-local VM and exposes a small discriminated state machine: running, suspended, failed, replay, or completed. The runtime retains a session only for a completed inline step when no attributes, hooks, waits, or replay divergence require the normal durable path.

On resume, the session verifies that every previously observed event is still an exact prefix, appends only the new events to the existing consumer, and wakes the suspended execution. A prefix mismatch or a second suspension boundary from an unobserved/discarded session moves it permanently to replay. Background workflow code can only enqueue in-memory work; the active observer remains the sole owner of durable writes and terminal drains.

This is intentionally invocation-local. Re-invocation, retry, rollover, queued execution, and process loss discard the session and use normal durable replay. The event log remains the source of truth.

Stacked on #2980, which prepares and caches replay payloads.

Validation

  • pnpm --filter @workflow/core typecheck
  • pnpm --filter @workflow/core build
  • all core tests: 70 files, 1,483 passed, 3 expected failures
  • retained-session coverage includes sequential suspensions, event-prefix divergence, late completion, discarded-session isolation, event-consumer quiescence, and telemetry context restoration
  • simplify and quality-code passes completed
  • final autoreview panel: Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh, zero findings

Benchmark

Run against workflow-server #632 at 063557ca5565d9dc3b182f602372c0222a7c9379. The benchmark-only SDK commit bed8acd4bc0042119c21dac2607f9c81f6e3b609 is the clean PR head d1907dc8fe22ddf331c7f047f6394e93224e9357 plus only the workflow-server preview URL override; the override is not present in this PR.

STSO windowCache only (#2980 + #632)Retained VM + cache (#2980 + #632)Average change
Steps 1-20286.6 ms166.6 ms-41.9%
Steps 101-120244.7 ms176.9 ms-27.7%
Steps 1001-1020556.4 ms228.3 ms-59.0%

In the exact 1,020-step trace, retained workflow.run calls are approximately 2.4 ms even at 3,000+ events. The remaining typical STSO is almost entirely the previous step_completed durable write plus the next step_started durable write. Two DynamoDB/write-path stalls dominate the retained late-window p90/p99; the VM remains constant during those samples.

The detailed trace decomposition and non-STSO metrics are in the benchmark comment below.

@vercel

vercelBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

@changeset-bot

changeset-botBot commented Jul 17, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 60ac8fb

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
NameType
@workflow/corePatch
workflowPatch
@workflow/buildersPatch
@workflow/cliPatch
@workflow/nextPatch
@workflow/nitroPatch
@workflow/vitestPatch
@workflow/web-sharedPatch
@workflow/webPatch
@workflow/world-testingPatch
@workflow/astroPatch
@workflow/nestPatch
@workflow/nuxtPatch
@workflow/rollupPatch
@workflow/sveltekitPatch
@workflow/vitePatch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

Summary

PassedFailedSkippedTotal
❌ ▲ Vercel Production1421322301683
✅ 💻 Local Development148302001683
✅ 📦 Local Production161702191836
✅ 🐘 Local Postgres161702191836
✅ 🪟 Windows15300153
✅ 📋 Other89401771071
❌ vercel-multi-region243027
Total72093510458289

❌ Failed Tests

▲ Vercel Production (32 failed)

nextjs-turbopack (29 failed):

  • DurableAgent e2e core basic text response
  • DurableAgent e2e core single tool call
  • DurableAgent e2e core multiple sequential tool calls
  • DurableAgent e2e core tool error recovery
  • DurableAgent e2e onStepFinish fires constructor + stream callbacks in order with step data
  • DurableAgent e2e onFinish fires constructor + stream callbacks in order with event data
  • DurableAgent e2e provider tools provider tool identity preserved across step boundaries
  • DurableAgent e2e provider tools mixed provider and function tools
  • DurableAgent e2e instructions string instructions are passed to the model
  • DurableAgent e2e timeout completes within timeout
  • DurableAgent e2e experimental_onStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onStepStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallStart (GAP) completes but callbacks are not called (GAP)
  • DurableAgent e2e experimental_onToolCallFinish (GAP) completes but callbacks are not called (GAP)
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • addTenWorkflow | wrun_41KXRJ0XDN0GYF1VG0Y3SC1DDP | 🔍 observability
  • wellKnownAgentWorkflow (.well-known/agent) | wrun_41KXRJ0NGM0GHVNQPYEJK1G3YD | 🔍 observability
  • promiseAllWorkflow | wrun_41KXRJ14110GV9A6P4ZZXKDJYH | 🔍 observability
  • promiseRaceWorkflow | wrun_41KXRJ16EC0GZTK8WVZ91846HT | 🔍 observability
  • promiseAnyWorkflow | wrun_41KXRJ1W0R0GJNBRANT4H32WTM | 🔍 observability
  • importedStepOnlyWorkflow | wrun_41KXRJ1VC60GJP6RWZFF8MSZRY | 🔍 observability
  • readableStreamWorkflow | wrun_41KXRJ23G50GQWBYHDW1QAW2XB | 🔍 observability
  • hookWorkflow | wrun_41KXRJ2HRY0GNG1066BP830Y64 | 🔍 observability
  • hookWorkflow is not resumable via public webhook endpoint | wrun_41KXRJ2TK70GVH4NXM9HF5JXZ5 | 🔍 observability
  • webhookWorkflow | wrun_41KXRJ2ZM80GGQERVSY0J41N0D | 🔍 observability
  • parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race | wrun_41KXRJ35JR0GMTCBR3T33CYHZQ | 🔍 observability
  • webhook route with invalid token
  • sleepingWorkflow | wrun_41KXRJ3Z6R0GS14DVVVXA18783 | 🔍 observability
  • parallelSleepWorkflow | wrun_41KXRJ4C0T0GM2Z3XJ6RYYJ9AB | 🔍 observability

nuxt (3 failed):

vercel-multi-region (3 failed)

nextjs-turbopack (3 failed):

  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: iad1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: sfo1 }) mints a tagged run ID and executes there
  • multi-region (world-vercel) explicit region: start({ region }) in the test process start({ region: fra1 }) mints a tagged run ID and executes there

Details by Category

❌ ▲ Vercel Production
AppPassedFailedSkipped
✅ astro126027
✅ example126027
✅ express126027
✅ fastify126027
✅ hono126027
❌ nextjs-turbopack121293
✅ nextjs-webpack15003
✅ nitro126027
❌ nuxt123327
✅ sveltekit14508
✅ vite126027
✅ 💻 Local Development
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 📦 Local Production
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🐘 Local Postgres
AppPassedFailedSkipped
✅ astro-stable128025
✅ express-stable128025
✅ fastify-stable128025
✅ hono-stable128025
✅ nextjs-turbopack-canary134019
✅ nextjs-turbopack-stable15300
✅ nextjs-webpack-canary134019
✅ nextjs-webpack-stable15300
✅ nitro-stable128025
✅ nuxt-stable128025
✅ sveltekit-stable14706
✅ vite-stable128025
✅ 🪟 Windows
AppPassedFailedSkipped
✅ nextjs-turbopack15300
✅ 📋 Other
AppPassedFailedSkipped
✅ e2e-local-dev-nest-stable128025
✅ e2e-local-dev-tanstack-start-128025
✅ e2e-local-postgres-nest-stable128025
✅ e2e-local-postgres-tanstack-start-128025
✅ e2e-local-prod-nest-stable128025
✅ e2e-local-prod-tanstack-start-128025
✅ e2e-vercel-prod-tanstack-start126027
❌ vercel-multi-region
AppPassedFailedSkipped
❌ nextjs-turbopack2430

📋 View full workflow run


Some E2E test jobs failed:

  • Vercel Prod: failure
  • Local Dev: failure
  • Local Prod: success
  • Local Postgres: success
  • Windows: success

Check the workflow run for details.

@NathanColosimo

NathanColosimo commented Jul 17, 2026

Copy link
Copy Markdown
ContributorAuthor

Retained VM benchmark and trace decomposition

Benchmark metrics

MetricCache onlyRetained VMObservation
STSO steps 1-20 avg / p75 / p90 / p99286.6 / 312 / 468 / 705 ms166.6 / 175 / 288 / 333 msavg -41.9%, p75 -43.9%
STSO steps 101-120 avg / p75 / p90 / p99244.7 / 262 / 304 / 382 ms176.9 / 161 / 227 / 666 msavg -27.7%, p75 -38.5%; one 501 ms completion write sets p99
STSO steps 1001-1020 avg / p75 / p90 / p99556.4 / 572 / 595 / 762 ms228.3 / 227 / 592 / 805 msavg -59.0%, p75 -60.3%; two durable-write stalls set p90/p99
Stream TTFS avg1,189.1 ms1,553.5 msno improvement; this starts before a retained resume and is noisy across single runs
Stream SL avg4,742.7 ms4,768.5 mseffectively unchanged (+0.5%)
Hook + stream TTFS avg1,453.8 ms1,988.9 msno improvement; same caveat as stream TTFS
Hook + stream SL avg4,937.2 ms5,028.3 mseffectively unchanged (+1.8%)

wo equals TTFS in these artifacts. The causal metric this PR changes is multi-step STSO after the first suspension; it does not optimize initial workflow/stream startup.

Does VM execution still scale with event count?

No, once the invocation retains the session:

  • 999 retained workflow.run spans: avg 2.63 ms, p50 2.34 ms, p75 2.44 ms, p90 2.63 ms
  • retained calls in the late 3,002-3,059-event range are mostly 1.28-2.95 ms
  • fresh replay at 0 events: 149.74 ms
  • fresh replay after invocation rollover at 2,348 events: 940.59 ms

There are only two fresh replays in this trace, so those are exact observations rather than a statistically useful replay distribution. The retained sample is approximately 1,020 calls. The result is still decisive for the scaling question: the 1,000th hot continuation does not replay 1,000 promises; it resumes the existing VM in about 2.4 ms.

Exact STSO decomposition

Each benchmark gap is approximately:

previous step_completed write + retained workflow.run + next step_started write + small harness/network residual

WindowArtifact STSO avgstep_completed avgretained workflow.run avgnext step_started avgSumResidual
Steps 1-20166.6 ms69.5 ms3.0 ms80.4 ms152.9 ms13.7 ms
Steps 101-120176.9 ms86.4 ms2.6 ms86.2 ms175.2 ms1.7 ms
Steps 1001-1020228.3 ms103.8 ms2.4 ms117.3 ms223.4 ms4.9 ms

Late-window outliers are visible directly in the spans:

  • gap 1010: step_completed 52.9 + resume 2.4 + step_started 531.7 = 586.9 ms; artifact STSO 591 ms
  • gap 1011: step_completed 727.1 + resume 2.5 + step_started 70.5 = 800.2 ms; artifact STSO 805 ms

The retained VM is not responsible for either tail sample.

One representative normal late step

Trace span 15082288459160472248 takes 129.44 ms:

  • step_started client/world create: 72.83 ms
  • workflow-server route: 53.29 ms
  • server materialization: 49.31 ms
  • DynamoDB query: 5.31 ms
  • DynamoDB transaction write: 25.33 ms
  • patch/insert work: 8.85 / 8.48 ms
  • local hydrate / user step / dehydrate: 0.29 / 0.03 / 0.34 ms
  • step_completed client/world create: 55.01 ms
  • workflow-server route: 35.80 ms
  • server materialization: 30.24 ms
  • bounded events query: 6.87 ms
  • retained workflow.run immediately afterward: approximately 2.4 ms

Conclusion

Retaining the VM removes the event-count-dependent replay cost for hot serial steps. It improves average STSO by 42%, 28%, and 59% in the three windows, and turns the late-window typical path from about 572 ms p75 into 227 ms p75.

It cannot reach 20 ms by itself. The remaining path contains two sequential durable transitions, each paying client/network, workflow-server, and DynamoDB materialization/write latency. The next performance change must collapse, overlap, or eliminate one or both round trips, for example by atomically completing the current step and claiming/starting the next one in one server transition. VM retention and that protocol optimization are complementary.

Alternatives and extension path

ApproachProsCons
Invocation-local retained VM (this PR)Minimal ownership change; event log stays authoritative; failure automatically falls back to replay; removes hot O(events) replayEnds at invocation/rollover; memory grows with the live session; does not remove durable writes
Keep an execution owner/actor across invocationsExtends the same constant-time resume across queue deliveries and can retain open streams/state longerRequires regional routing, leases, fencing, failover, eviction, and a precise durability protocol
Snapshot/restore the VMCould survive process movement without replaying source-level workflow historyNode vm execution state and pending promises are not serializable; needs a different isolate/runtime or engine-specific snapshots and is substantially more complex
Fuse step_completed and next step_startedDirectly attacks the approximately 220 ms that remains after VM retention; one atomic server transition can remove a round tripChanges World/server contracts and retry/idempotency semantics; needs careful handling for parallel steps, hooks, and failures
Keep stateless replay and optimize bundle/hydration cachesSimplest durability and horizontal scaling model; useful for cold starts regardlessCannot remove the O(events) workflow execution/replay path, so it does not solve late hot STSO alone

The chosen state machine is deliberately invocation-local so it can later sit inside an actor/owner without changing its session API. The most valuable immediate follow-up is protocol fusion; cross-invocation ownership is useful only after deciding that retaining locality is worth its operational cost.

Review and verification

  • core typecheck and build pass
  • core suite: 70 files, 1,483 passed, 3 expected failures
  • final simplify/quality pass completed
  • final autoreview used Codex gpt-5.6-sol xhigh and Claude Opus 4.8 xhigh; both returned zero findings

@github-actions

github-actionsBot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 60ac8fb · Fri, 17 Jul 2026 17:39:10 GMT · run logs

Backend: vercel · app: nextjs-turbopack

MetricScenarioAvg (ms)P75 (ms)P90 (ms)P99 (ms)Samples
TTFSstream1256 (-0.8%)1660 🔴1708 🔴1853 🔴30
TTFShook + stream1625 (+22%)1916 🔴1992 🔴2064 🔴30
STSO1020 steps (1-20)176 (-41%)192 🔴278 🔴293 🔴19
STSO1020 steps (101-120)151 (-51%)156 🔴188 🔴221 🔴19
STSO1020 steps (1001-1020)171 (-73%)176 🔴270 🔴290 🔴19
WOstream1256 (-0.8%)16601708185330
WOhook + stream1625 (+22%)19161992206430
SLstream3881 (+288%)5720 🔴5832 🔴5920 🔴30
SLhook + stream4571 (+134%)5550 🔴5635 🔴5800 🔴30

Avg deltas compare against the most recent benchmark run on main at the time of this run.

Metrics — TTFS: time to first step body execution · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (time outside step bodies, client start → last step body exit) · SL: stream latency (first chunk write → visible to the reader)

Scenarios — stream: one step that streams chunks back to the client; no hooks, so the run stays in turbo mode · hook + stream: registers a hook before the same streaming step, which exits turbo mode · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges

🟢/🔴 mark percentiles within/above target. Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · STSO (1-20) 20/30/60 · STSO (101-120) 30/45/90 · STSO (1001-1020) 40/60/120

TTFS/WO compare client vs deployment clocks and SL compares the step runner’s clock vs the client’s (NTP-synced in CI). WO ends at the last step body exit, the closest observable proxy for the final step-completion request.

@NathanColosimo

Copy link
Copy Markdown
ContributorAuthor

Overriden by #3046 and #3047

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@NathanColosimo