JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Re-enable acceleration of Vector512<long>.op_Multiply - #111832

Merged
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq
Jan 26, 2025
Merged

JIT: Re-enable acceleration of Vector512<long>.op_Multiply#111832
EgorBo merged 1 commit into
dotnet:mainfrom
saucecontrol:mullq

Conversation

@saucecontrol

Copy link
Copy Markdown
Member

This was a regression in 9.0, from #103555

https://godbolt.org/z/11hs3Kqdd

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jan 25, 2025
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jan 25, 2025
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

// Emulate NI_AVX512DQ_VL_MultiplyLow with SSE41 for SIMD16
}
else
else if (simdSize != 64)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, I think this needs Avx512DQ ISA check

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're treating any of the Vector512 methods other than IsSupported as intrinsic, that implies we have the full baseline AVX-512 set (F,DQ,BW,CD,VL). It's a bit confusing because some of the import paths assert or check that, but most don't. I'm actually cleaning up some of those redundant asserts in a different branch now.

@EgorBoEgorBoJan 25, 2025

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah ok, I thought that Vector512.IsHardwareAccelerated only relies on AVX512F, but looks like DOTNET_EnableAVX512DQ=0 turns it off so it's ok.

@saucecontrolsaucecontrolJan 26, 2025

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, IsHardwareAccelerated 😄

The rule is Vector512.IsHardwareAccelerated will return false unless all of the following are satisfied:

  1. Baseline AVX-512 support (F,CD,DW,BW,VL)
  2. Not throttling based on CPUID check for Skylake-X and others with severe downclocking for 512-bit vector instructions, or DOTNET_PreferredVectorBitWidth: >= 512
  3. No DOTNET_PreferredVectorBitWidth: < 512

The rule for whether Vector512 methods actually import as intrinsic is only that we have the baseline AVX-512 set, meaning IsHardwareAccelerated may return false, but all methods may actually be accelerated anyway.

So the fact that we're importing the methods for Vector512 as intrinsic in the first place means the ISA requirements have already been met.

Vector128 and Vector256 are a bit different, because the baseline ISA requirement may not be enough to accelerate all methods.

Vector256.IsHardwareAccelerated returns true only if AVX2 is supported, but we attempt to import methods as intrinsic as long as AVX is supported. Since many of the methods require AVX2 for acceleration, they have an extra check for AVX2 and then fall back to managed if it's not available. Hence all the (simdSize != 32) || compOpportunisticallyDependsOn(InstructionSet_AVX2) checks.

Similar checks are not included for Vector128, because the base requirement is SSE2, so almost all methods can be accelerated, minus a few that require SSE4.1 and check for it explicitly.

Clear as mud, I know...

@EgorBo
EgorBo merged commit d5c8265 into dotnet:mainJan 26, 2025
grendello added a commit to grendello/runtime that referenced this pull request Jan 27, 2025
* main: (22 commits)
Clean up Stopwatch a bit (dotnet#111834)
JIT: Fix embedded broadcast simd size (dotnet#111638)
Revert potential UB due to aliasing + more WB removals (dotnet#111733)
re-enable acceleration of Vector512<long>.op_Multiply (dotnet#111832)
Handle OSSL 3.4 change to SAN:othername formatting
JIT: Fix stack allocated arrays for NativeAOT (dotnet#111827)
JIT: enhance RBO inference for similar compares to constants (dotnet#111766)
JIT: Don't run optSetBlockWeights when we have PGO data (dotnet#111764)
[Android] Make sure RuntimeFlavor=CoreCLR when clr subset is specified (dotnet#111821)
Change empty subject test certificate to include a critical SAN.
Fix reversed code offsets in GcInfo (dotnet#111792)
Swap some libraries areas between leads (dotnet#111816)
Add left-handed spherical and cylindrical billboards (dotnet#109605)
JIT: revise `optRelopImpliesRelop` to always set `reverseSense` (dotnet#111803)
Fix Zip64ExtraField handling (dotnet#111802)
Add build support for Android+CoreCLR (dotnet#110471)
arm64: Add bic(s) compact encoding (dotnet#111452)
JIT: Ensure `BBF_PROF_WEIGHT` flag is set when we have PGO data (dotnet#111780)
Add support for AVX10.2, Add AVX10.2 API surface and template tests (dotnet#111209)
JIT: Preliminary for enabling inlining late devirted calls (dotnet#111782)
...
@saucecontrol
saucecontrol deleted the mullq branch January 28, 2025 01:22
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Feb 27, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@saucecontrol@EgorBo