JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Accelerate more casts on x86 - #116805

Closed
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2
Closed

JIT: Accelerate more casts on x86#116805
saucecontrol wants to merge 14 commits into
dotnet:mainfrom
saucecontrol:lng2flt2

Conversation

@saucecontrol

@saucecontrolsaucecontrol commented Jun 19, 2025

Copy link
Copy Markdown
Member

This is the follow-up to #114597. It addresses the invalid removal of intermediate double casts in long->double->float chains and does more cleanup/normalization in fgMorphExpandCast, removing more of the platform-specific logic.

It also adds some more acceleration to x86. With this, all integral<->floating casts are accelerated with AVX-512, and all 32-bit integral<->floating casts are accelerated with the baseline ISAs.

Example AVX-512 implementation of floating->long/ulong casts:

 G_M38981_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref, nogc <-- Prolog IG
- sub esp, 8- vzeroupper - ;; size=6 bbWeight=1 PerfScore 1.25+ ;; size=0 bbWeight=1 PerfScore 0.00
G_M38981_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}, byref
- vmovsd xmm0, qword ptr [esp+0x0C]- sub esp, 8- ; npt arg push 0- ; npt arg push 1- vmovsd qword ptr [esp], xmm0- call CORINFO_HELP_DBL2LNG- ; gcr arg pop 2- ;; size=19 bbWeight=1 PerfScore 6.25+ vmovsd xmm0, qword ptr [esp+0x04]+ vcmpsd k1, xmm0, xmm0, 7+ vcvttpd2qq xmm1 {k1}{z}, xmm0+ vcmpsd k1, xmm0, qword ptr [@RWD00], 29+ vpcmpeqd xmm0, xmm0, xmm0+ vpsrlq xmm1 {k1}, xmm0, 1+ vmovd eax, xmm1+ vpextrd edx, xmm1, 1+ ;; size=51 bbWeight=1 PerfScore 18.50
G_M38981_IG03: ; bbWeight=1, epilog, nogc, extend
- add esp, 8
ret 8
- ;; size=6 bbWeight=1 PerfScore 2.25+ ;; size=3 bbWeight=1 PerfScore 2.00+RWD00 dq	43E0000000000000h

Example SSE2/SSE4.1 implementation of uint->floating casts:

 G_M24329_IG01: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000000 {}
;; size=3 bbWeight=1 PerfScore 0.25
G_M24329_IG02: ; bbWeight=1, gcrefRegs=00000000 {}, byrefRegs=00000002 {ecx}, byref
; byrRegs +[ecx]
- push 0- ; npt arg push 0- push dword ptr [ecx]- ; npt arg push 1- call [CORINFO_HELP_LNG2DBL]- ; byrRegs -[ecx]- ; gcr arg pop 2- fstp qword ptr [esp]- movsd xmm0, qword ptr [esp]+ mov eax, dword ptr [ecx]+ xorps xmm0, xmm0+ cvtsi2sd xmm0, eax+ movaps xmm1, xmm0+ addsd xmm1, qword ptr [reloc @RWD00]+ movaps xmm2, xmm0+ blendvpd xmm0, xmm1+ movsd qword ptr [esp], xmm0
fld qword ptr [esp]
- ;; size=21 bbWeight=1 PerfScore 10.00+ ;; size=36 bbWeight=1 PerfScore 15.83
G_M24329_IG03: ; bbWeight=1, epilog, nogc, extend
add esp, 8
ret ;; size=4 bbWeight=1 PerfScore 1.25
+RWD00 dq	41F0000000000000h

Full diffs

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 19, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 19, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@saucecontrol
saucecontrol marked this pull request as ready for review June 19, 2025 23:16
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

cc @tannergooding

Comment threadsrc/coreclr/jit/hwintrinsiccodegenxarch.cpp
Comment threadsrc/coreclr/jit/morph.cpp
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
Comment threadsrc/coreclr/jit/decomposelongs.cpp
Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
@saucecontrol

Copy link
Copy Markdown
MemberAuthor

The codegen for AVX-512 floating->long regressed after merging in main (likely something with #116983). I'll give it another look.

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

#117485 resolved the regression. Diffs look good again.

Comment threadsrc/coreclr/jit/decomposelongs.cpp Outdated
CorInfoType srcBaseType = CORINFO_TYPE_UNDEF;
CorInfoType dstBaseType = CORINFO_TYPE_UNDEF;

if (varTypeIsFloating(srcType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we know if the compiler is CSEing this check with the above check as expected?

-- Asking since manually caching might be a way to win some throughput back.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not for sure, but I generally assume C++ compilers will handle 'obvious' ones like this. It should be noted, the throughput hit to x86 directly correlates with the number of casts that are now inlined.

i.e. the only significant throughput hit is on the coreclr_tests collection

image

which is also the one that had the most casts in it

image

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍. The biggest concern is the TP hit to minopts. It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think #117512 will reduce the hit a bit.

It may be desirable to leave that using the helper there so that floating-point heavy code doesn't start up slower.

This is interesting, because the same argument could apply to the complicated saturating logic that we have for x64 as well. #97529 introduced a similar throughput regression, and although it was done for correctness instead of perf, the throughput hit could have been avoided by using the helper in minopts there too.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the throughput hit could have been avoided by using the helper in minopts there too.

AFAIR, the JIT throughput hit there ended up being very minimal (and often an improvement). It was the perf score and code output size that regressed, which was expected.

If there was a significant perf score hit to minopts, then yes the same would apply here and it would likely be beneficial to ensure that is doing the "better" thing as well.

Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
Comment threadsrc/coreclr/jit/flowgraph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

@saucecontrol there's some merge conflicts on this one and it is also dependent on #117571 going in first, correct?

@saucecontrol

Copy link
Copy Markdown
MemberAuthor

My plan was to trim this down to just the new acceleration for x86, but yeah, that would build on the part split into #117571. I'll set this to draft in the meantime.

@saucecontrol
saucecontrol marked this pull request as draft August 8, 2025 16:37
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@tannergooding