[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[Wasm RyuJit] Enable native wasm fast tail calls - #129134

Merged
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls
Jun 15, 2026
Merged

[Wasm RyuJit] Enable native wasm fast tail calls#129134
AndyAyersMS merged 4 commits into
dotnet:mainfrom
AndyAyersMS:wasm-fast-tailcalls

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Jun 8, 2026

Copy link
Copy Markdown
Member

Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call / return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to undo the prolog adjustment, so the callee receives the incoming shadow-stack pointer.

Leaving implicit tail calls as normal calls for now.

Set FEATURE_FASTTAILCALL=1 and FEATURE_TAILCALL_OPT=1. Fast tail calls
lower to return_call / return_call_indirect. Tag the SP arg so codegen
adds compLclFrameSize to undo the prolog adjustment, so the callee
receives the incoming shadow-stack pointer.
CopilotAI review requested due to automatic review settings June 8, 2026 18:59
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 8, 2026
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@kg PTAL
fyi @dotnet/wasm-contrib @dotnet/jit-contrib

Passes various Pri-0 tail call tests. We emit ~4K tail calls in SPC.

Using an LIR flag may raise some hackles; happy to consider alternatives.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Enables WebAssembly fast tail calls in CoreCLR RyuJIT and wires up shadow-stack/SP handling so wasm return_call / return_call_indirect can be emitted correctly.

Changes:

  • Turn on FEATURE_FASTTAILCALL and FEATURE_TAILCALL_OPT for TARGET_WASM.
  • Tag the wasm shadow-stack/SP argument for fast tail calls in RA and adjust it in codegen to undo the prolog’s SP delta.
  • Relax a fast-tailcall eligibility check that is stack-based and not applicable to wasm’s local-based argument passing.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Show a summary per file
FileDescription
src/coreclr/jit/targetwasm.hEnables fast tail calls + opportunistic tail calls for wasm.
src/coreclr/jit/regallocwasm.cppTags the well-known wasm shadow-stack pointer arg for fast tail calls.
src/coreclr/jit/morph.cppSkips an arg-stack-space constraint that doesn’t apply to wasm.
src/coreclr/jit/lir.hAdds a wasm-specific LIR flag to mark the fast-tailcall SP arg.
src/coreclr/jit/codegenwasm.cppEmits INS_end for tailcall “jmp epilog” blocks and adjusts SP arg / return type handling for tail calls.

Comment threadsrc/coreclr/jit/targetwasm.h
Comment threadsrc/coreclr/jit/regallocwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@SingleAccretionSingleAccretion left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

What's the benefit of using guaranteed-tailcall return_call for implicit tailcalls?

I was imagining we could use implicit tailcalls for shadow stack only since it has some benefits w.r.t. zero-sized shadow frames.

Not sure what to make of this ... are you saying the underlying engine can do this instead in most cases? Or that's not worth doing in general?

@SingleAccretion

SingleAccretion commented Jun 8, 2026

Copy link
Copy Markdown
Contributor

Or that's not worth doing in general?

return_call is the WASM equivalent of .NET .tail. It constrains the final code generator for the benefit of predictable semantics. Fast tailcalls are about performance, so this kind of change should come with some (measured) performance benefit. I don't know whether in the current engines there will be such a benefit. In principle, constraining the code generator should be strictly worse.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

@JulieLeeMSFTJulieLeeMSFT added the arch-wasm WebAssembly architecture label Jun 9, 2026
@SingleAccretion

Copy link
Copy Markdown
Contributor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I suggest we just enable all the implicit cases for now to get test coverage and then come back later to figure out the impact on perf.

It's not just performance though, it has impacts in terms of debugging and such. I guess I am not really understanding the underlying motivation. Test coverage for return_call would only be important for explicit tail calls i. e. .tail. Is it important to have fully working .tail on WASM currently?

No, it's not important.

@dotnet/wasm-contrib any thoughts here? Should we leave implicit disabled for now?

@AndyAyersMS

AndyAyersMS commented Jun 11, 2026

Copy link
Copy Markdown
MemberAuthor

This recent paper A Snapshot of the Performance of
Wasm Backends for Managed Languages
claims tail call performance is actually fairly decent (see section 4.2), at least in Bigloo Scheme. So at least there's some evidence in favor...

Thanks @davidwrighton for finding this.

@SingleAccretion

SingleAccretion commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

So at least there's some evidence in favor...

It seems expected that tailcalls in general are an improvement (we wouldn't have implicit tailcalls [on native targets] otherwise). My point above is that return_call is really like .tail and not like call; ret (implicitly tailcallable in IL). From my point of view, it would not make sense for a language compiling to IL to sort of randomly insert .tail for every tail-positioned call, even if on our current Jit implementation it would happen to produce better code in some cases.

Keep FEATURE_FASTTAILCALL=1 so explicit `.tail` calls still lower to
return_call, but set FEATURE_TAILCALL_OPT=0 so we don't yet opt in to
implicit tail calls while we shake out the new wasm fast-tailcall path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@AndyAyersMSAndyAyersMS mentioned this pull request Jun 12, 2026
16 tasks
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

I'll disable implicit tail calls for now. Added a note to #121865 to reconsider later.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@adamperlin can you review?

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated

@adamperlinadamperlin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand the difference in signature logic for tail calls. Otherwise, I don't necessarily mind the LIR::WasmFastTailCallSP flag, though it could be cleaner to do an ADD of the framesize back to the SP in IR as Single is suggesting.

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
AndyAyersMSand others added 2 commits June 12, 2026 16:40
Use call->gtReturnType + call->gtRetClsHnd to compute the wasm-level
result type uniformly for both regular and fast tail calls, instead
of forking on params.isJump. For fast tail calls fgCanFastTailCall
already ensures caller/callee result types are compatible.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Express the fast tail-call SP adjustment in IR as ADD(SP, FRAME_SIZE),
where GT_FRAME_SIZE is a new wasm-only LIR leaf that codegen resolves
to the (post-regalloc) compLclFrameSize value.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings June 12, 2026 23:42

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@adamperlin

Copy link
Copy Markdown
Contributor

This looks good to me! The CI failures appear to be infra related or known issues.

@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

/ba-g unrelated failures (this PR is R2R wasm specific)

@AndyAyersMS
AndyAyersMS merged commit 9907f43 into dotnet:mainJun 15, 2026
132 of 139 checks passed
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview6 milestone Jun 17, 2026
eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Set FEATURE_FASTTAILCALL=1. Explicit tail calls lower to return_call /
return_call_indirect. Tag the SP arg so codegen adds compLclFrameSize to
undo the prolog adjustment, so the callee receives the incoming
shadow-stack pointer.
Leaving implicit tail calls as normal calls for now.
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 18, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@AndyAyersMS@SingleAccretion@adamperlin@JulieLeeMSFT