Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by felixweinberger · Pull Request #2547 · modelcontextprotocol/typescript-sdk · GitHub
Skip to content

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) - #2547

Merged
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle
Jul 27, 2026
Merged

Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)#2547
felixweinberger merged 7 commits into
v1.xfrom
fweinberger/sse-keepalive-lifecycle

Conversation

@felixweinberger

@felixweinbergerfelixweinberger commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Follow-up to #2538. Keep-alive timers are now owned by the stream they were armed for instead of being tracked in a transport-level map keyed by stream id: startKeepAlive returns the handle, the stream's own cancel/cleanup clears it, and the interval clears itself on write failure. The shared timer map, stopKeepAlive, and the close() timer sweep are removed. Cancel callbacks only delete the stream mapping when it still points at their own stream (matching main), a resume closes the superseded stream instead of leaving it hanging, the POST path arms after its fallible awaits, and non-finite keepAliveMs values disable keep-alive (values above 2^31-1 are clamped).

Motivation and Context

The standalone GET stream and its resumed successors share one stream id, so with the shared timer map a few ordinary disconnect/reconnect orderings tear down the wrong timer:

  • After a client reconnects and resumes, the stale connection's cancel callback stops the resumed stream's keep-alive and deletes its stream mapping. The reconnecting client silently stops receiving frames and server notifications — the disconnect loop from SSE stream disconnected: TypeError: terminated #1211 comes back for exactly the clients keep-alive is meant to protect.
  • In the POST path the timer is armed before await writePrimingEvent(...). If eventStore.storeEvent rejects, the handler returns 400 and the Response is discarded, so nothing can ever cancel the stream; the timer keeps firing until close().
  • close() during a replayEventsAfter await doesn't stop the continuation from arming a timer afterwards, which nothing can clear.
  • keepAliveMs: NaN / Infinity / values above 2^31-1 pass the <= 0 guard, and setInterval clamps such delays to ~1ms — flooding every SSE stream with comment frames. Number(process.env.SOME_UNSET_VAR) is an easy way to hit this.

How Has This Been Tested?

Seven new regression tests (fake timers), each failing on current v1.x before the fix. Full test suite passes. Also verified against a real http.Server over raw sockets: after the stale connection drops, the resumed stream keeps receiving keep-alive frames and server notifications, and keepAliveMs: NaN writes nothing.

Breaking Changes

None.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Update (second commit): the review comments on the first commit pointed at the same stale-capture class inside send(). Rather than spot-guards, the second commit gives send() a consistent model: the event store is the source of truth, live streams are best-effort delivery.

  • Response events are stored even when the request stream is disconnected (matching the standalone stream path), so a polling client can always retrieve the response via Last-Event-ID.
  • After the store write, the stream registration is re-read, and events the resumed stream's replay already delivered are skipped (tracked per registration), preventing lost or duplicated responses when a resume races an in-flight write. The standalone notification path gets the same guard.
  • Request correlations are released once a response is safely replayable; without an event store, or in JSON response mode, completing against a missing stream still surfaces an error.
  • Session initialization and stream registration are also guarded against a transport that closed while the request body or the onsessioninitialized callback was pending.

Seven more regression tests, each pinned red-green.

Keep-alive timers were tracked in a transport-level map keyed by stream
id. The standalone GET stream and its resumed successors share one id,
so a stale connection's cancel callback could stop the resumed stream's
keep-alive and delete its stream mapping, and the interval's error
handler could clear the wrong timer.
Timers are now owned by the stream they were armed for: startKeepAlive
returns the handle, the stream's own cancel/cleanup clears it, and the
interval clears itself on write failure. Cancel callbacks only delete
the stream mapping when it still points at their own stream. A resume
closes the superseded stream cleanly, and a resume that completes after
the transport closed (or after its stream was cancelled) returns 404
instead of registering a stream nothing can clean up.
The POST path now arms keep-alive after its fallible awaits and
releases the stream and request correlations if they fail. Non-finite
keepAliveMs values disable keep-alive, and values above 2^31-1 are
clamped instead of firing every millisecond.
@felixweinberger
felixweinberger requested a review from a team as a code ownerJuly 24, 2026 15:42
@changeset-bot

changeset-botBot commented Jul 24, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 0f983d0

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@pkg-pr-new

pkg-pr-newBot commented Jul 24, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@modelcontextprotocol/sdk@2547

commit: 0f983d0

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@mattzcarey

Copy link
Copy Markdown
Contributor

Hmm yeh this is better :) thanks!!

A response completing while its SSE stream was being resumed could be
written to the evicted stream (lost) or delivered twice when the event
store made it replay-visible before the write resolved. A stream that
disconnected before its response completed also left the response
unstored, so a polling client could never retrieve it, and the request
correlation maps leaked.
send() now stores response events even when the request stream is
disconnected (matching the standalone stream path), re-reads the stream
registration after the store write, skips events the resumed stream's
replay already delivered, and releases request correlations once the
response is safely replayable. Without an event store, or in JSON
response mode where a replay can never settle the pending response,
completing against a missing stream still surfaces an error.
Also guards session initialization and stream registration against a
transport that closed while the request body or the session
initialization callback was pending.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
felixweinbergerand others added 2 commits July 27, 2026 14:05
Extract keep-alive validation into a shared sseKeepAlive helper matching
main: non-finite and sub-millisecond intervals disable keep-alive, and
oversized delays are clamped instead of firing every millisecond.
Also ported from the main-branch implementation: handleRequest returns
404 once the transport is closed instead of registering streams nothing
can clean up, SSE responses send X-Accel-Buffering: no (and the POST
stream gains no-transform) so proxy buffering does not swallow frames,
and a DELETE request closes the transport even when the onsessionclosed
callback throws.
send() now also releases request correlations before reporting an
undeliverable response, so the entries don't outlive it.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment thread.changeset/sse-keepalive-timer-lifecycle.md Outdated
send() released a completed response's correlations to Last-Event-ID
replay whenever an event store was configured, but a client can only
resume a stream if it actually holds a cursor: clients that never
received an id-bearing event (for example pre-2025-11-25 clients whose
stream carried no notifications) can never issue the resume, so their
lost responses were released silently where they previously surfaced
through onerror.
Request streams are now marked resumable when an id-bearing event is
delivered on them - the priming event, a stored notification, or a
replay - and the silent hand-off applies only to those streams. A
response completing while the transport closes is a no-op instead of a
spurious error, failed POST registrations release their tracking entry,
and the changeset now also describes the send() delivery-model changes.

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, two candidates were examined and ruled out this pass: the sseKeepAlive.ts camelCase filename (matches the existing convention of its siblings, e.g. webStandardStreamableHttp.ts, despite CLAUDE.md's hyphens rule), and the resume supersede path applying to a half-open predecessor — when the event store implements getStreamIdForEventId, the 409 conflict check runs first, so the supersede-without-conflict-check shape only exists for stores without that optional method.

Extended reasoning...

This run's inline comments (two nits, one pre-existing InMemoryEventStore issue) are posted separately; this note only records what else was examined and refuted so a later pass doesn't re-explore it from scratch. No prior run left a ruled-out note on this PR, so this is the first and only such record. It is informational, not a correctness guarantee.

Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
The closed-transport no-op in send() ran after the response was recorded
in _requestResponseMap, so a send parked on the event-store write when
close() swept the maps re-populated them and stranded the response
payload and its correlations on the closed transport. Release them
before returning.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Repeated close() re-ran the stream sweep and re-fired onclose. The main
branch guards close() re-entry; this restores the same guard, which the
DELETE handler's unconditional close in its finally block now relies on.
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
Comment threadsrc/server/webStandardStreamableHttp.ts
@felixweinberger
felixweinberger merged commit e3f3daa into v1.xJul 27, 2026
12 checks passed
@felixweinberger
felixweinberger deleted the fweinberger/sse-keepalive-lifecycle branch July 27, 2026 17:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v1Issues / PRs related to v1.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@felixweinberger@mattzcarey