Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Guard NetworkTransport connect continuation against stale reconnectio… - #178

Open
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation
Open

Guard NetworkTransport connect continuation against stale reconnectio…#178
algal wants to merge 1 commit into
modelcontextprotocol:mainfrom
algal:fix-networktransport-connect-continuation

Conversation

@algal

Copy link
Copy Markdown

This PR fixes a crash in NetworkTransport.connect() where a delayed reconnection task could resume a stale continuation after a new connect attempt started, causing a crash.

Motivation and Context

The bug occurs when a delayed reconnection task from a previous connect attempt runs after a new connect attempt has started.

While using iMCP.app, I saw that app crash when a client (imcp-server) started and then died quickly. Crash logs showed CheckedContinuation.resume(throwing:) from NetworkTransport.handleReconnection. What was happening was that when NetworkTransport.connect() is called, it uses a checked continuation and schedules reconnection work with backoff on failure/cancellation. If a reconnect attempt is scheduled and a new connect() call happens before that delayed task fires, the old task can still resume the previous continuation. That double‑resumes a CheckedContinuation and traps (EXC_BREAKPOINT / SIGTRAP).

How Has This Been Tested?

Trying this with the iMCP.app is the most easy way to repro the crash, and what I used to ensure that the fix removed the crash. You just run iMCP, then try running its imcp-server and CTRL-C it immmediately a few times. Eventually, this brings down the iMCP process as well.

I have not produced a test which reproduces the bug in isolation, outside of iMCP, but I can do so if that would help. The sequence would work as follows

Repro (before fix):

  1. Create a NetworkTransport (default reconnection config).
  2. Call connect(), then cancel the underlying connection quickly enough to enter handleReconnection.
  3. Call connect() again before the reconnection backoff delay elapses.
  4. When the delayed task fires, it resumes the old continuation and traps (EXC_BREAKPOINT / SIGTRAP).

I found that swift test fails on Swift 6.2.3 due to existing strict-concurrency errors in Tests/MCPTests/ClientTests.swift.

How this fix works

The continuation‑resume guard was a single boolean reset on each connect() call. Delayed reconnection tasks from a previous connect attempt could still run after a new connect had reset the guard, allowing a second resume on the old continuation.

This fix works as follows:

  • Introduce a per‑connect UUID (connectContinuationID).
  • Thread that ID through handleConnectionReady/Failed/Cancelled and handleReconnection.
  • Resume only if the ID matches the current connect attempt, and invalidate it after resuming.
  • Ignore delayed reconnection tasks from stale attempts.

Breaking Changes

None.

Besides the fix, behavior is unchanged; this only prevents stale tasks from resolving an old continuation.

This is an internal-only change; the only observable difference is that stale reconnection tasks no longer resolve a previous connect() attempt.

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Checklist

  • I have read the MCP Documentation
  • My code follows the repository's style guidelines
  • New and existing tests pass locally
  • I have added appropriate error handling
  • I have added or updated documentation as needed

Additional context on use of AI

I (algal) have reviewed this PR, performed the before/after tests above, and I believe the analysis I have stated above is correct. I also have a background in Swift and of native development. However, I am not deeply familiar with this code base, and I did rely on AI in the developing this fix and drafting this PR text.

I know there's a lot of slop PRs flying around these days, so I wanted to be transparent about that. I am happy to be corrected or steered

@DePasqualeOrg

Copy link
Copy Markdown

This issue, along with many others, has already been resolved in my fork, which @movetz, @stallent, and I are discussing merging into this repository.

Instead of using a checked continuation with callback-based state handling (which requires the UUID-based guard proposed in this PR), my fork uses AsyncStream to bridge NWConnection state updates:

privatefunc waitForConnectionReady()asyncthrows{letstateStream= AsyncStream<NWConnection.State>{ continuation in
connection.stateUpdateHandler ={ state in
continuation.yield(state)switch state {case.ready,.failed,.cancelled:
continuation.finish()default:break}}}forawaitstatein stateStream {switch state {case.ready:returncase.failed(let error):throw error
case.cancelled:throwMCPError.internalError("Connection cancelled")
// ...
}}}

This approach inherently avoids the double-resume race condition because:

  • A fresh stream is created for each connect() call.
  • The stream finishes on terminal states, leaving no lingering continuation.
  • My fork removes in-connect reconnection scheduling entirely — on failure, it throws immediately rather than scheduling a delayed retry. Reconnection is handled in receiveLoop() for post-connect failures.

@algal

Copy link
Copy Markdown
Author

@DePasqualeOrg Cool! It sounds like your solution is better. I hope it is merged. :)

@DePasqualeOrg

DePasqualeOrg commented Jan 20, 2026

Copy link
Copy Markdown

It won't be, unless it's replicated by someone else, because the new maintainers of this package want to take a patchwork approach to improvements. My full-featured fork will continue to be developed separately here: DePasqualeOrg/swift-mcp

@movetz
movetz requested a review from stallentJanuary 27, 2026 20:17
@movetzmovetz added the bug Something isn't working label Jan 27, 2026
@toasterbook88

Copy link
Copy Markdown

Maintainer triage on February 24, 2026: this transport fix looks valuable but currently conflicts with main. Please rebase onto current main and rerun CI for a decision.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bugSomething isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@algal@DePasqualeOrg@toasterbook88@movetz