JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen - #108750

Merged
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601
Aug 7, 2025
Merged

JIT: Refine assert for RMW check in crc32 and hw intrinsic codegen#108750
jakobbotsch merged 9 commits into
dotnet:mainfrom
jakobbotsch:fix-108601

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Oct 10, 2024

Copy link
Copy Markdown
Member

The crc32 HW intrinsic is a 2 operand node that produces a register; however, the underlying instruction is an RMW instruction that destructively modifies op1. For that reason the codegen needs to issue a move of op1 into the target register before crc32 is issued.

There is an assert that this mov does not overwrite op2 in the process. When LSRA builds uses for the operands it needs to mark op2 as delay freed to ensure op2 and the target do not end up in the same register. However, due to optimizations LSRA may skip this marking when op1 is a last use. Thus we can end up hitting the assert.

The most logical solution seems to be to actually ensure that op2 is always delay freed, such that it never shares the register with the target. However, the handling for other RMW nodes (like GT_SUB) does not have similar logic to this; instead it relies on the situation to not happen except for one particular case where it is ok: when op1 and op2 are actually the same local. This PR enhances the assert to match the checks that happen for other RMW nodes.

Also enhance a similar assert for arm64 HW intrinsics.

Fix#108601
Fix#117612

The crc32 HW intrinsic is a 2 operand node that produces a register;
however, the underlying instruction is an RMW instruction that
destructively modifies `op1`. For that reason the codegen needs to issue
a move of op1 into the target register before `crc32` is issued.
There is an assert that this `mov` does not overwrite op2 in the
process. When LSRA builds uses for the operands it needs to mark op2 as
delay freed to ensure op2 and the target do not end up in the same
register. However, due to optimizations LSRA may skip this marking when
op1 is a last use. Thus we can end up hitting the assert.
The most logical solution seems to be to actually ensure that op2 is
always delay freed, such that it never shares the register with the
target. However, the handling for other RMW nodes (like `GT_SUB`) does
not have similar logic to this; instead it relies on the situation to
not happen except for one particular case where it is ok: when `op1` and
`op2` are actually the same local. This PR enhances the assert to match
the checks that happen for other RMW nodes.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Oct 10, 2024
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

This matches the logic in genCodeForBinary from here:

#ifdef DEBUG
unsigned lclNum1 = (unsigned)-1;
unsigned lclNum2 = (unsigned)-2;
GenTree* op1Skip = op1->gtSkipReloadOrCopy();
GenTree* op2Skip = op2->gtSkipReloadOrCopy();
if (op1Skip->OperIsLocalRead())
{
lclNum1 = op1Skip->AsLclVarCommon()->GetLclNum();
}
if (op2Skip->OperIsLocalRead())
{
lclNum2 = op2Skip->AsLclVarCommon()->GetLclNum();
}
assert(GenTree::OperIsCommutative(oper) || (lclNum1 == lclNum2));
#endif

To be honest I do not see the constraint properly encoded in LSRA for the common RMW nodes or for crc32, which is a bit scary. In other words, if we have IR like

/- op1 LCL_VARV01 last use
/- op2 LCL_VARV02GT_SUB

then I think op2 will have a use created that is not marked delay freed. That means the following register assignment is perfectly allowable:

/- op1 LCL_VARV01 last use rdx
/- op2 LCL_VARV02 rcx
GT_SUB rcx

however, this assignment would hit the assert in genCodeForBinary. The only reason this does not happen seems to be due to preferencing making it so that LSRA ends up not picking this assignment.

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp Outdated
CopilotAI review requested due to automatic review settings August 6, 2025 15:17

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR refines an assertion in the CRC32 hardware intrinsic code generation to handle edge cases where LSRA optimizations may cause operand registers to overlap with the target register. The fix aligns the CRC32 assert with the handling used for other RMW (read-modify-write) operations.

Key changes:

  • Enhanced the assertion in CRC32 codegen to allow the case where both operands are the same local variable
  • Extracted and reused common logic for checking if two trees represent the same local variable
  • Simplified existing duplicate assertion logic in binary operation codegen

Reviewed Changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
src/coreclr/jit/hwintrinsiccodegenxarch.cppEnhanced CRC32 assertion to include same-local-variable check
src/coreclr/jit/codegenxarch.cppAdded genIsSameLocalVar helper function and simplified binary operation assertion
src/coreclr/jit/codegen.hAdded declaration for the new genIsSameLocalVar helper function

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@jakobbotschjakobbotsch changed the title JIT: Refine assert for RMW check in crc32 codegenJIT: Refine assert for RMW check in crc32 and hw intrinsic codegenAug 6, 2025
jakobbotschand others added 2 commits August 6, 2025 17:23
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@tannergooding can you please take another look at this?

Also cc @dotnet/jit-contrib and PTAL @amanasifkhalid

Comment threadsrc/coreclr/jit/hwintrinsiccodegenarm64.cpp
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Out of space infra failure

@jakobbotsch
jakobbotsch merged commit c993ba2 into dotnet:mainAug 7, 2025
103 of 105 checks passed
@jakobbotsch
jakobbotsch deleted the fix-108601 branch August 7, 2025 14:19
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 7, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

4 participants

@jakobbotsch@tannergooding@amanasifkhalid