Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode by madzarm · Pull Request #611 · modelcontextprotocol/rust-sdk · GitHub
Skip to content

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode - #611

Merged
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition
Jan 13, 2026
Merged

fix: use Semaphore instead of Notify in OneshotTransport to prevent race condition during stateless mode#611
alexhancock merged 1 commit into
modelcontextprotocol:mainfrom
madzarm:fix/stateless-transport-race-condition

Conversation

@madzarm

@madzarmmadzarm commented Jan 9, 2026

Copy link
Copy Markdown
Contributor

Motivation and Context

Fixes issue #610

OneshotTransport uses tokio::sync::Notify to signal when the transport should close. However, Notify::notify_waiters() only wakes futures that are currently awaiting - if the notification fires before receive() starts waiting, the signal is lost forever.

This creates a race condition in serve_inner() where tokio::select! can drop the receive() future between iterations:

  1. receive() returns the initial message
  2. send() completes with Response/Error, calls notify_waiters()
  3. select! picks join_next() branch, drops the receive() future
  4. Next iteration: receive() creates NEW notified() future
  5. This future waits forever - the notification already fired
  6. Serve loop hangs, stream never closes

In stateless HTTP mode (Lambda, Cloud Functions), this causes requests to hang until timeout even though the tool executed successfully.

Solution

Replace Notify with Semaphore. Unlike notifications, permits persist until acquired:

  • add_permits(1) stores the permit even if no one is waiting
  • acquire().await returns immediately if a permit exists, or waits if not
  • Dropping an acquire() future doesn't consume the permit (cancel-safe)

From tokio docs:

"This method is cancel safe. If acquire is used as the event in a tokio::select! statement and some other branch completes first, then no permits are acquired."

How Has This Been Tested?

  • Deployed to AWS Lambda with API Gateway in stateless mode
  • Load tested with concurrent requests during cold starts
  • Specifically tested error paths (downstream 404s) which previously triggered the race ~3-5% of the time
  • After fix: zero hangs across thousands of requests

Breaking Changes

None. This is an internal implementation change. The OneshotTransport public API is unchanged.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context

The race window is small but hits reliably under load. In our production environment, ~3-5% of cold start requests would hang, always on error responses (possibly due to faster code paths when is_error: true).

The fix is minimal - only changes the synchronization primitive from Notify to Semaphore with equivalent semantics but proper cancellation safety.

@github-actionsgithub-actionsBot added the T-core Core library changes label Jan 9, 2026

@alexhancockalexhancock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice catch and fix, thank you

@madzarm
madzarmforce-pushed the fix/stateless-transport-race-condition branch from e8d7c73 to 2e657baCompareJanuary 13, 2026 21:05
@madzarm

Copy link
Copy Markdown
ContributorAuthor

Hi @alexhancock thanks for the quick response! I've shortened the commit message to fix the commit lint failure. I'd appreciate it if you could approve the workflow, thank you!

@alexhancock
alexhancock merged commit c4a6829 into modelcontextprotocol:mainJan 13, 2026
11 checks passed
@alexhancockalexhancock mentioned this pull request Jan 14, 2026
takumi-earth pushed a commit to earthlings-dev/rmcp that referenced this pull request Jan 27, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

T-coreCore library changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@madzarm@alexhancock@scutuatua-crypto