Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Implement wasm codegen for sub-16 SIMD load/store by tannergooding · Pull Request #131000 · dotnet/runtime · GitHub
Skip to content

Implement wasm codegen for sub-16 SIMD load/store - #131000

Merged
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store
Jul 18, 2026
Merged

Implement wasm codegen for sub-16 SIMD load/store#131000
tannergooding merged 4 commits into
dotnet:mainfrom
tannergooding:wasm-sub16-simd-load-store

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so they live as a v128 with the low 8/12 bytes populated. Previously the wasm JIT bailed out via the ins_Load/ins_Store NYIs for these types; this implements the split lane load/store sequences instead.

The emitted sequences are:

  • simd8 load: v128.load64_zero 0
  • simd8 store: v128.store64_lane 0, lane 0
  • simd12 load: local.get addr; v128.load64_zero 0; v128.load32_lane 8, lane 2
  • simd12 store: tee the value into a v128 temporary, v128.store64_lane 0, lane 0 for the low 8 bytes, then re-materialize the address and v128.store32_lane 8, lane 2 for the upper 4 bytes

v128.load64_zero fills lanes 0-1 (zeroing the rest); the trailing lane store/load handles bytes 8-11 for the Vector3 case.


The TYP_SIMD12 address is forced multiply-used (loads and heap stores re-materialize it for the trailing lane op). The local-to-stack store rewrite (RewriteLocalStackStore) produces a STOREIND(LCL_ADDR, value) whose address is a re-materializable GT_LCL_ADDR, so that case is excluded from multiply-use and codegen re-emits the frame pointer directly. Because that synthesized STOREIND is not revisited by the main collection walk, its internal v128 tee register is requested in RewriteLocalStackStore.


Measured on the corelib crossgen2 browser SuperPMI collection (27,540 contexts): hard asserts drop from 403 to 30, eliminating 373SIMD8/SIMD12 load/store asserts with no regressions. The residual 30 are the pre-existing NYIRAW oper catch-all, unrelated to this change.

Full effectiveness requires #130866 (which removes the shadowing SIMD-ABI bailouts for SIMD params/locals/stores/call-args); standalone, this change already clears the 373 asserts above on that collection.

Note

This PR description was drafted with the assistance of GitHub Copilot.

Vector2 (TYP_SIMD8) and Vector3 (TYP_SIMD12) have no native wasm valtype, so
they live as a v128 with the low 8/12 bytes populated. Emit the split lane
loads/stores instead of bailing out via the ins_Load/ins_Store NYIs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 21:45
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 17, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
10 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the WASM JIT’s SIMD codegen to handle sub-16-byte SIMD memory operations by emitting lane-based load/store sequences for TYP_SIMD8 (Vector2) and TYP_SIMD12 (Vector3), and updates lowering/regalloc to ensure required multi-use addressing and internal v128 temporaries are available.

Changes:

  • Add SIMD8 load support via v128.load64_zero and route SIMD12 loads/stores through new split-lane helper sequences in WASM codegen.
  • Force multiply-use of SIMD12 addresses in lowering where codegen needs to re-materialize/re-push the address.
  • Ensure regalloc requests an internal v128 temp for SIMD12 stores (including the local-to-stack store rewrite path).
Show a summary per file
FileDescription
src/coreclr/jit/regallocwasm.cppRequests an internal v128 register for SIMD12 store indirections (including rewrite-introduced STOREIND).
src/coreclr/jit/lowerwasm.cppForces multiply-use of SIMD12 indir addresses when codegen needs to reuse/re-materialize them.
src/coreclr/jit/instr.cppImplements ins_Load(TYP_SIMD8) as INS_v128_load64_zero for WASM.
src/coreclr/jit/codegenwasm.cppImplements split-lane load/store sequences for SIMD12 and lane store for SIMD8 storeind; adds helper routines.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 2

Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Casting the frame offset to unsigned before passing it as a wasm memarg
defeated the emitter's offset >= 0 invariant, silently turning an unexpected
negative offset into a large positive one. Compute a signed offset and
noway_assert it, matching emitter::emitIns_S.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:04
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0 new

CopilotAI review requested due to automatic review settings July 17, 2026 22:14

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 4

Comment threadsrc/coreclr/jit/codegenwasm.cpp
Comment threadsrc/coreclr/jit/codegenwasm.cpp Outdated
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Comment threadsrc/coreclr/jit/lowerwasm.cpp
Assert the wasm frame address is FP-based before re-emitting the frame pointer
in the simd12 lane helpers, matching the block-copy codegen. Update the
multiply-use DEBUGARG reasons since simd12 now forces it regardless of faulting.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 17, 2026 22:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 1

Comment threadsrc/coreclr/jit/codegenwasm.cpp
@lewing

Copy link
Copy Markdown
Member

End-to-end validation on the wasm R2R prototype tree (which actually runs wasm R2R crossgen, complementary to the SuperPMI assert numbers):

Stacked this on #130866 with a matched crossgen2 + wasm cross-jit and crossgen'd the built JIT/Regression SIMD tests with no punt flags:

  • 9/23 → 21/23 previously-failing SIMD tests now crossgen clean. All 12 sub-16 SIMD8/SIMD12 load/store cases clear.
  • Emitted sequences confirmed: Vector2v128.load64_zero / store64_lane; Vector3load64_zero + load32_lane (and the tee'd store64_lane + store32_lane).

The 2 still-failing are unrelated to this PR: one harness/reference artifact, and one nested try/catch-with-filter EH case (Runtime_129972.M()) that asserts in fgwasm.cppBBF_CATCH_RESUMPTION — separate wasm EH work.

LGTM from the end-to-end side.

Note

This comment (and its validation run) was generated with GitHub Copilot.

@tannergooding
tannergooding enabled auto-merge (squash) July 18, 2026 00:40

@lewinglewing left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@tannergooding
tannergooding merged commit 0def554 into dotnet:mainJul 18, 2026
140 checks passed
@tannergooding
tannergooding deleted the wasm-sub16-simd-load-store branch July 18, 2026 16:18
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 19, 2026
lewing added a commit to lewing/runtime that referenced this pull request Jul 21, 2026
…Shuffle (Adam's shuffle-mask fix)
Fixes the invalid i8x16.shuffle immediate (out-of-range lane) that made the
SIMD composite fail wasm validation. Resolved 2 conflicts vs dotnet#130866/dotnet#131000:
- codegenwasm.cpp: dropped redundant ins local (superseded by dotnet#131000 store restructure)
- lowerwasm.cpp: took shuffle case + IsCnsVec() immediate handling from dotnet#130991
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@tannergooding@lewing@adamperlin