Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Fix HWIntrinsic codegen for elided scalar/vector reinterprets - #131155

Merged
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen
Jul 21, 2026
Merged

Fix HWIntrinsic codegen for elided scalar/vector reinterprets#131155
tannergooding merged 1 commit into
dotnet:mainfrom
tannergooding:tannergooding-fix-hwintrinsic-reinterpret-codegen

Conversation

@tannergooding

Copy link
Copy Markdown
Member

Fixes#131137.

PR #130444 started removing transparent scalar/vector reinterpret HWINTRINSIC nodes (CreateScalarUnsafe, GetLower/GetLower128, ToVector256Unsafe/ToVector512Unsafe) during lowering when the consumer is another HWINTRINSIC. That's correct in general -- the consumer reads the value from a register at its own size -- but two x64 codegen sites keyed their decision off the operand's post-lowering type, which the elision changes. Both now produce wrong code.


ConvertToVector128Int* / ConvertToVector256Int* (the reported crash)

These have a vector overload (Vector128<T>) and a pointer overload (T*), both lowering to pmovzx*. Codegen picked between them with varTypeIsSIMD(op1). Once an elided CreateScalarUnsafe leaves the vector overload's operand scalar-typed, that proxy misfires and codegen takes the memory-load path -- reading the scalar value as if it were an address (the reported NullReferenceException; a checked JIT asserts in emitxarch.cpp). Fixed by selecting the overload from the stable node->OperIsMemoryLoad() metadata (aux-type driven), matching the generic table path already used elsewhere in the file.


AVX2 gather VSIB index width

A gather selects its VSIB index width (xmm vs ymm) from the index operand's own width (indexOp->TypeIs(TYP_SIMD32)). An elided GetLower/ToVector*Unsafe on the index changes that width, so the wrong VEX.L is encoded and the hardware reads the wrong number of indices (e.g. a vpgatherqd with a GetLower()-narrowed index gathered 4 elements instead of 2 -- a silent wrong result, not visible in the JIT disasm since it always prints the index as xmm). The index width can't be recovered in codegen, so this is fixed in lowering: the reinterpret elision is skipped when the node is a gather's index operand, since that width is load-bearing.


I also audited the rest of hwintrinsiccodegenxarch.cpp, the store-containment paths in codegenxarch.cpp, and emitxarch.cpp: every other size/attr decision derives from node metadata (node->TypeGet(), GetSimdSize(), GetSimdBaseType()), and the operandSize >= expectedSize check in IsContainableHWIntrinsicOp prevents a narrowed operand from being contained undersized. These two sites were the only operand-width-driven exceptions.

Added Runtime_131137 covering both fixed paths (Sse41/Avx2 convert-from-scalar and the gather-with-narrowed-index case); it fails without the fix and passes with it.

Note

This PR was authored with the assistance of GitHub Copilot.

After dotnet#130444 began removing transparent CreateScalarUnsafe/GetLower/
ToVector*Unsafe nodes during lowering, two x64 codegen sites that keyed off
an operand's post-lowering type produced wrong code:
* ConvertToVector128Int*/ConvertToVector256Int* selected the pointer
(memory-load) overload via varTypeIsSIMD(op1); an elided CreateScalarUnsafe
left the vector overload's operand scalar-typed, so it loaded from the value
as if it were an address. Use node->OperIsMemoryLoad() instead.
* An AVX2 gather selects its VSIB index width (xmm vs ymm) from the index
operand's own width; an elided GetLower/ToVector*Unsafe on the index changed
that width and gathered the wrong number of elements. Skip the elision when
the node is a gather's index operand, since that width is load-bearing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings July 21, 2026 15:37
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 5 pipeline(s).
11 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

@dotnet/jit-contrib, @EgorBo, @dhartglassMSFT

A fix for #131137. I did an audit of the other code paths and found one additional failure that I also added handling for.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes two x64 RyuJIT HWIntrinsic codegen bugs caused by lowering eliding “transparent” scalar/vector reinterpret nodes, where a couple of codegen decisions incorrectly depended on the post-lowering operand type instead of stable node metadata.

Changes:

  • Update xarch lowering to preserve width-changing reinterprets when they are used as the AVX2 gather index operand, preventing VSIB width (xmm vs ymm) from being mis-encoded.
  • Update xarch HWIntrinsic codegen to select vector vs pointer overloads for the ConvertToVector{128,256}Int* intrinsics using node->OperIsMemoryLoad() rather than varTypeIsSIMD(op1->TypeGet()).
  • Add a JitBlue regression test covering both the convert-from-scalar and gather-with-narrowed-index scenarios.
Show a summary per file
FileDescription
src/coreclr/jit/lowerxarch.cppSkips elision of GetLower/ToVector*Unsafe-style reinterprets when the user is an AVX2 gather and the node is the index operand, preserving VSIB width selection.
src/coreclr/jit/hwintrinsiccodegenxarch.cppUses OperIsMemoryLoad() (aux-type-driven) to choose the correct convert overload path after reinterpret elision changes operand TypeGet().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csAdds xUnit-based regression coverage for Sse41/Avx2 convert-from-scalar and gather with index.GetLower().
src/tests/JIT/Regression/JitBlue/Runtime_131137/Runtime_131137.csprojAdds the test project for the new JitBlue regression scenario.

Copilot's findings

  • Files reviewed: 4/4 changed files
  • Comments generated: 0

@JulieLeeMSFTJulieLeeMSFT added the Priority:1 Work that is critical for the release, but we could probably ship without label Jul 21, 2026
@tannergooding
tannergooding merged commit 8aab0ce into dotnet:mainJul 21, 2026
149 of 151 checks passed
@tannergooding
tannergooding deleted the tannergooding-fix-hwintrinsic-reinterpret-codegen branch July 21, 2026 22:05
tannergooding added a commit that referenced this pull request Jul 22, 2026
Follow up to
#131155 (comment).
`Runtime_131137` was added with a standalone `.csproj` and
`RequiresProcessIsolation=true`, but the test meets none of the
isolation rules in
[`requiresprocessisolation.md`](https://github.com/dotnet/runtime/blob/main/docs/workflow/testing/coreclr/requiresprocessisolation.md)
-- it sets no environment variables, no host config, and no process-wide
state. It''s just a `[ConditionalFact]` in a `Runtime_131137` namespace
with no custom `Main`, so it merges cleanly.
This removes the standalone project and adds the source into the shared
`Regression_ro_2.csproj` runner, matching the rest of the regression
tests.
----------
I also audited the other recently-added `JitBlue` tests that still carry
standalone csprojs. All of them are justified: the immediate neighbors
`Runtime_130844`/`130845`/`130846` set `CLRTestEnvironmentVariable`, the
remaining ISO ones set
`CLRTestEnvironmentVariable`/`CLRTestTargetUnsupported`, and
`Runtime_8980` disables the xunit wrapper generator. `Runtime_131137`
was the only one with no trigger.
CC. @EgorBo
> [!NOTE]
> This PR description was drafted by Copilot.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 23, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:1Work that is critical for the release, but we could probably ship without

Projects

None yet

4 participants

@tannergooding@EgorBo@JulieLeeMSFT