Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Improve the handling of SIMD comparisons - #104944

Merged
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality
Jul 19, 2024
Merged

Improve the handling of SIMD comparisons#104944
tannergooding merged 2 commits into
dotnet:mainfrom
tannergooding:simd_equality

Conversation

@tannergooding

Copy link
Copy Markdown
Member

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 16, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@tannergooding
tannergooding marked this pull request as ready for review July 16, 2024 21:07
@tannergooding

Copy link
Copy Markdown
MemberAuthor

Got some good diffs (https://dev.azure.com/dnceng-public/public/_build/results?buildId=743416&view=ms.vss-build-web.run-extensions-tab) for Arm64 and x64.

windows arm64

Overall (-3,548 bytes)
MinOpts (-36 bytes)
FullOpts (-3,512 bytes)

windows x64

Overall (-11,220 bytes)
MinOpts (-1,610 bytes)
FullOpts (-9,610 bytes)

For arm64 this is just from allowing op_Equality and op_Inequality to fold. For example:

- mvni v16.4s, #0- mvni v17.4s, #0- cmeq v16.4s, v16.4s, v17.4s- uminp v16.4s, v16.4s, v16.4s- umov x0, v16.d[0]- cmn x0, #1- cset x0, eq- ;; size=28 bbWeight=1 PerfScore 5.00+ mov w0, #1

For x64 we get all the same scenarios, but we also further optimize comparisons against AllBitsSet. For example:

 vpcmpeqd xmm1, xmm1, xmm1
- vpcmpd k1, xmm1, xmm0, 4- kortestb k1, k1- sete al+ vptest xmm0, xmm1+ setb al

@xtqqczze

Copy link
Copy Markdown
Contributor

@MihuBot -dependsOn 104944

MihuBot/runtime-utils#537

Doesn't seem to have helped with regressions in #104488 (compare with MihuBot/runtime-utils#519).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The "regresssions" in that PR are minor size regressions, they aren't necessarily perf regressions.

I plan on handling those separately as its a bit more complex to solve given it requires looking across two levels. What's happening is that or is used by and and the and is used in op_Equality. We only see that or is used by and so we decide to fold it to ternarylogic.

It's a tradeoff in a minor size increase vs additional complexity in the JIT, where the minor size increase in this edge case is representative of significant size reduction in other cases.


vpor is 1 cycle; 0.33 reciprocal throughput. vptest is <4 cycles; 1 cycle reciprocal throughput

vmovaps for register to register is 0-1 cycle (typically 0 as its handled by the register renamer on most CPUs, vpternlogd is 1 cycle; 0.33 reciprocal throughput, and vptest` remains the same

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

@xtqqczze

xtqqczze commented Jul 17, 2024

Copy link
Copy Markdown
Contributor

So it's essentially a 1-to-1 trade here that isn't as meaningful to "fix" immediately

In case, these are not diffs from this PR.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

this diff is 1.47:1 (on latency, single iteration)

You're putting a lot of reliance on mca, which itself isn't a fully accurate tool and is giving a best potential estimate of the performance for a very particular microarchitecture. It's not necessarily representative of actual execution performance, latencies, or other behaviors of real world application code.

We do not microtune to that level in the BCL because it is not relevant to most real world workloads (and is far too microarchitecture dependent). We instead find a reasonable balance between overall readability, maintainability, and performance across the entire application. That may include trading a couple cycles in one piece of code to win back cycles in a lot of other more common code.

The specific case you've called out is something that will be addressed, it is not relevant to block the PR you've linked on and is not related to the improvements that are being done in this PR. It's simply additional improvements that could be had on top to ensure that a specific case of load bottlenecked vpternlog can be worse than two independent bitwise operations.

case NI_Vector512_op_Equality:
#endif // !TARGET_ARM64 && !TARGET_XARCH
{
if (varTypeIsFloating(simdBaseType))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how is this path different from case NI_Vector128_op_Equality: above? if it's only for floats, then why it's not an assert?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, it's when only one of the operand is constant?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it’s for when one is all nan, which is an optimization we can do for float/double

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, although, the last time I tried to constant fold vector comparisons I had no diffs, so it's very likely we have no test coverage there (I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

case NI_Vector64_op_Equality:
#elif defined(TARGET_XARCH)
case NI_Vector256_op_Equality:
case NI_Vector512_op_Equality:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is also NI_VectorX_EqualsAll (unless they're normalized to op_Equality somewhere). Btw the last time I tried to constant fold these, you told me that is odd to cover only EQ/NE relation operators ;-) #85584 (comment)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We’ve already added support for all the other comparisons (elementwise eq/ge/gt/le/lt/ne), what was remaining was the ==/!= operators, which this pr covers.

EqualsAll/Any and the other All/Any APIs are then imported as elementwise compare + op ==/!=, so this covers the full set

@tannergooding

Copy link
Copy Markdown
MemberAuthor

(I presume your diffs aren't from <cns_vec> ==/!= <cns_vec>)

That’s what most of the diffs are from, we have a lot of coverage now, especially since vector2/3/4 are in managed and matrix4x4 was accelerated. There’s also some cases in other test/code where we’re getting opts from the switch to use more xplat APIs and centralized helpers

@tannergooding
tannergooding merged commit e0ecd1f into dotnet:mainJul 19, 2024
@tannergooding
tannergooding deleted the simd_equality branch July 19, 2024 18:55
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 19, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tannergooding@xtqqczze@EgorBo