Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Wasm irreducible loop transformation by AndyAyersMS · Pull Request #121728 · dotnet/runtime · GitHub
Skip to content

Wasm irreducible loop transformation - #121728

Merged
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation
Nov 22, 2025
Merged

Wasm irreducible loop transformation#121728
AndyAyersMS merged 30 commits into
dotnet:mainfrom
AndyAyersMS:WasmIrreducibleLoopTransformation

Conversation

@AndyAyersMS

@AndyAyersMSAndyAyersMS commented Nov 18, 2025

Copy link
Copy Markdown
Member

If the Wasm DFS detects improper loop headers, then we have irreducible loops that cannot be expressed in Wasm control flow.

To fix this, run a pass to find the SCCs in the flow graph using Kosaraju's algorithm. Then invoke this algorithm recursively on the subgraph formed from the nodes in each SCC, minus the SCC entry nodes (nodes in the SCC with preds not in the SCC). Repeat until all "nested" SCCs are identified. This represents the full set of irreducible loops we need to transform. Note no SCCs share headers but nested SCCs will share interior blocks.

Single-entry SCCs are reducible loops and don't require any special processing as they can be emitted as Wasm lops. But multi-entry SCCs are irreducible loops and must be transformed.

So we transform each multi-emtry SCC (working inner to outer) by creating a per-SCC control var and dispatch block. Each SCC header is assigned an index from 0...N-1, where N is the number of headers in that SCC. The dispatch block switches to each the headers based on their index and the control var. Each pre-existing edge to the header is then logically split and the index var is assigned the index for that header and retargeted to the dispatch node. As an optimization and to handle some unsplittable edges, if an SCC header's pred has the header as its only successor, we put the control var assignment into the pred instead of splitting the edge.

This transforms each multi-entry SCC into a single-entry reducible loop. In checked builds we verify by rerunning the DFS and assert that there are no longer any improper headers.

Note there are other strategies for resolving SCCs into reducible loops that might offer better performance; we are intentionally picking something simple.

Defer handling cases where the original DFS found non-funclet blocks that could only be reached via EH, as we do not yet have a way of describing how Wasm control can reach such blocks. We will revisit this once we have the Wasm EH model design in place. Such cases are fairly rare (eg a try/catch that ends with a goto or return).

We currently run the SCC transform before lower to allow lower the chance to optimize the switch and because we introduce new IR. There is a risk that a sufficiently clever later phase (say one that could do block cloning or jump threading) might undo the dispatch structure and recreate an irreducible loop, but that doesn't seem to happen. The subsequent Wasm control flow phase will also assert that its run of Wasm DFS does not have any improper headers.

Continuation of #120534.

Contributes to #121178.

AndyAyersMSand others added 24 commits November 6, 2025 11:34
Determine how to emit Wasm control flow from the JIT's control flow graph.
Relies on loop-aware RPO to determine the block order. Currently only
handles the main method. Assumes irreducible loops have been fixed
upstream (which is not yet guaranteed; bails out if not so).
Doesn't actually do any emission, just prints a textual description in
the JIT dump (along with a dot markup version).
Uses only LOOP and BLOCK. Tries to limit the extent of BLOCK.
Run for now as an optional phase even if not targeting Wasm, to
do some stress testing.
Contributes to dotnet#121178
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Loops for Wasm control flow codegen don't involve EH or runtime mediated
control flow transfers.
Implement a custom block successor enumerator for Wasm, and adjust `fgRunDFS`
to allow using this and also to generalize how the DFS is initiated. Use
this to build a "Wasm" DFS. In that DFS handle both the main method and
all funclets (by specifying funclet entries as additional DFS starting points).
Update the loop finding code to make suitable changes when it is driven from
a "Wasm" DFS instead of the typical all successor / all predecessor DFS.
Remove the restriction in the Wasm control flow codegen that only handles
the main method; now it works for the main method and all funclets.
Contributes to dotnet#121178.
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Nov 18, 2025
@am11am11 added the arch-wasm WebAssembly architecture label Nov 18, 2025
@AndyAyersMS
AndyAyersMS marked this pull request as ready for review November 18, 2025 16:32
CopilotAI review requested due to automatic review settings November 18, 2025 16:32
Comment threadsrc/coreclr/jit/fgwasm.cpp Outdated
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.cpp
Comment threadsrc/coreclr/jit/fgwasm.h
Comment threadsrc/coreclr/jit/fgwasm.h Outdated
Comment threadsrc/coreclr/jit/jiteh.cpp
@kg

kg commented Nov 19, 2025

Copy link
Copy Markdown
Contributor

The parts I understand LGTM

Comment threadsrc/coreclr/jit/fgwasm.cpp
@AndyAyersMSAndyAyersMS mentioned this pull request Nov 21, 2025
@AndyAyersMS

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib any other comments?

If not, I need one of you to approve this.

kg
kg approved these changes Nov 22, 2025
@AndyAyersMS
AndyAyersMS merged commit b1d5443 into dotnet:mainNov 22, 2025
110 of 112 checks passed
Comment on lines +4943 to +4952
#ifdef DEBUG
// If we are going to simulate generating wasm control flow,
// transform any strongly connected components into reducible flow.
//
if (JitConfig.JitWasmControlFlow() > 0)
{
DoPhase(this, PHASE_DFS_BLOCKS_WASM, &Compiler::fgDfsBlocksAndRemove);
DoPhase(this, PHASE_WASM_TRANSFORM_SCCS, &Compiler::fgWasmTransformSccs);
}
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it intentional this was placed before PHASE_ASYNC, even though it is eventually going to have to be after?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not intentional. It should be moved to just after the async transformation.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973

Comment on lines +273 to +277
if (BitVecOps::IsMember(m_traits, m_blocks, pred->bbPostorderNum))
{
// Pred is in the scc, so not an entry edge
continue;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that we see predecessors here that won't be in the DFS tree such that pred->bbPostorderNum is something undefined? Should this be guarded on that?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We run a DFS+Remove pass just before, but being defensive here can't hurt. Let me look into this.

Comment on lines +463 to +483
// Dump subgraph as dot
{
JITDUMP("digraph scc_%u_nested_subgraph%u {\n", m_num, nestedCount);
BitVecOps::Iter iterator(m_traits, nestedBlocks);
unsigned int poNum;
bool first = true;
while (iterator.NextElem(&poNum))
{
BasicBlock* const block = m_dfsTree->GetPostOrder(poNum);

JITDUMP(FMT_BB ";\n", block->bbNum);

WasmSuccessorEnumerator successors(m_comp, block, /* useProfile */ true);
for (BasicBlock* const succ : successors)
{
JITDUMP(FMT_BB " -> " FMT_BB ";\n", block->bbNum, succ->bbNum);
}
}

JITDUMP("}\n");
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Put this under #ifdef DEBUG and if (verbose)?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in #121973.

Comment on lines +750 to +752
// TODO: if we had a BV iter that worked from highest set
// bit to lowest, we could iterate the subset directly
// and avoid searching here.

@jakobbotschjakobbotschNov 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can use BitVecOps::VisitBitsReverse for this.

In fact I would suggest switching away from the BV iter in most places here and unify the interface of Scc with FlowGraphNaturalLoop.

Comment on lines +767 to +774
if (sccs.Height() > 0)
{
for (int i = 0; i < sccs.Height(); i++)
{
Scc* const scc = sccs.Bottom(i);
scc->Finalize();
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The emptiness check looks unnecessary.

//
void FgWasm::WasmFindSccsCore(BitVec& subset, ArrayStack<Scc*>& sccs, BasicBlock** postorder, unsigned postorderCount)
{
SccMap map(Comp()->getAllocator(CMK_WasmSccTransform));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this map be sparse, or could it just be a flat map indexed by the postorder indices, since that mapping is dense?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Initially this will need to cover the entire method, so I think we can create a flat array for that and then reuse it for subsequent subset cases.

Comment on lines +363 to +366
for (BasicBlock* const pred : block->PredBlocks())
{
advance();
hasPred = true;
if (!BitVecOps::IsMember(&m_traits, subgraph, pred->bbPostorderNum))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similarly, does this need to be guarded on m_dfsTree->Contains(pred)?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 26, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@AndyAyersMS@kg@jakobbotsch@adamperlin@am11