Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add support for utilizing F16C instructions on xarch - #127094

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd
Apr 18, 2026
Merged

Add support for utilizing F16C instructions on xarch#127094
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:half-simd

Conversation

@tannergooding

@tannergoodingtannergooding commented Apr 17, 2026

Copy link
Copy Markdown
Member

Since #122649 had to be reverted due to the ABI concerns, this is a simpler initial change that works with the existing ABI and on hardware with AVX2 support (not just AVX512-FP16 capable hardware).

This should provide a nice win across most existing hardware and we can follow up with a PR that does similar for the AVX512-FP16 instructions that allow directly accelerated arithmetic operations, rather than only handling conversions.

Before

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M16314_IG01: ;; offset=0x0000 4883EC28 subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M16314_IG02: ;; offset=0x0004 0FB7C9 movzxrcx,cx FF156BA74500 call[System.Half:op_Explicit(System.Half):float]90nop ;; size=10 bbWeight=1 PerfScore 3.50G_M16314_IG03: ;; offset=0x000E 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code: 19; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M32250_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M32250_IG02: ;; offset=0x0000 FF2572A74500 tail.jmp[System.Half:op_Explicit(float):System.Half] ;; size=6 bbWeight=1 PerfScore 2.00; Total bytes of code: 6

After

; Method Program:HalfToSingle(System.Half):float (FullOpts)G_M15861_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15861_IG02: ;; offset=0x0000 0FB7C1 movzxrax,cx C5F96EC0 vmovd xmm0,eax C4E27913C0 vcvtph2psxmm0,xmm0 ;; size=12 bbWeight=1 PerfScore 6.25G_M15861_IG03: ;; offset=0x000C C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 13; Method Program:SingleToHalf(float):System.Half (FullOpts)G_M15413_IG01: ;; offset=0x0000 ;; size=0 bbWeight=1 PerfScore 0.00G_M15413_IG02: ;; offset=0x0000 C4E3791DC000 vcvtps2phxmm0,xmm0,0 C5F97EC0 vmovd eax,xmm0 0FB7C0 movzxrax,ax ;; size=13 bbWeight=1 PerfScore 6.25G_M15413_IG03: ;; offset=0x000D C3 ret ;; size=1 bbWeight=1 PerfScore 1.00; Total bytes of code: 14

CopilotAI review requested due to automatic review settings April 17, 2026 21:36
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 17, 2026
@tannergooding

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel -amd

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;publicclassBenchmarks{staticvoidMain(string[]args){BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args);}floatdataF32=float.Pi;HalfdataF16=Half.Pi;[Benchmark]publicfloatHalfToSingle()=>(float)dataF16;[Benchmark]publicHalfSingleToHalf()=>(Half)dataF32;}

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds initial xarch JIT support to accelerate System.Halffloat explicit conversions by recognizing Half.op_Explicit as a named intrinsic and lowering it to AVX2/F16C conversion instructions where available, without changing the existing ABI.

Changes:

  • Mark System.Half and the Half(float) / float(Half) explicit operators as [Intrinsic] so the JIT can recognize them.
  • Add a new named intrinsic (NI_System_Half_op_Explicit) and importer expansion that emits AVX2 conversion HW intrinsics (vcvtps2ph / vcvtph2ps) for xarch.
  • Update xarch HW intrinsic lists + containment/perf metadata to support the new conversion instructions.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
src/libraries/System.Private.CoreLib/src/System/Half.csMarks Half and key explicit operators as [Intrinsic] to enable JIT recognition.
src/coreclr/jit/namedintrinsiclist.hAdds a named intrinsic ID for System.Half.op_Explicit.
src/coreclr/jit/importercalls.cppRecognizes Half.op_Explicit and expands it to AVX2 conversion intrinsic sequences on xarch.
src/coreclr/jit/importer.cppAdds helper routines to pack/unpack scalar Half values through SIMD nodes.
src/coreclr/jit/compiler.hDeclares helper routines and adds isSystemHalfClass type recognition.
src/coreclr/jit/hwintrinsiclistxarch.hAdds AVX2 conversion intrinsics for half<->single vector conversions.
src/coreclr/jit/lowerxarch.cppExtends containment logic to support the new conversion/store patterns.
src/coreclr/jit/emitxarch.cppAdds perf characteristics entries for the new conversion instructions.

Comment threadsrc/coreclr/jit/importer.cpp Outdated
Comment threadsrc/coreclr/jit/importercalls.cpp
Comment threadsrc/coreclr/jit/importercalls.cpp
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, @EgorBo, @kg for review

@dotnet/intel and @jkotas as an FYI on the alternative approach. AVX512-FP16 support can be done nearly identically, it's just a bigger PR. I'll pull the changes from #122649 after this lands. We can then look at the ABI handling and ensuring Half is properly passed in a floating-point register in the future.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Benchmark (EgorBot/Benchmarks#132) is too small to get good results...

The realistic is that HalfToSingle is roughly 4.16x faster and SingleToHalf is about 1.76x faster. Changing from about 28 instructions w/ 3 memory accesses and 25 instructions w/ 0 memory accesses, respectively, to about 2 instructions with 0 memory accesses.

kg
kg approved these changes Apr 18, 2026
@tannergooding
tannergooding enabled auto-merge (squash) April 18, 2026 03:48
@tannergooding
tannergooding merged commit 29bafa7 into dotnet:mainApr 18, 2026
173 of 185 checks passed
@tannergooding
tannergooding deleted the half-simd branch April 18, 2026 14:59
tannergooding added a commit that referenced this pull request Apr 19, 2026
Follow up to #127094
The instruction was added a while ago but only started being utilized in
the above PR. However, the instruction needs to be `MR` rather than `RM`
encoded or it can cause invalid results in some scenarios.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 19, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@kg