Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use popcount intrinsincs in BitOperations - #85944

Closed
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt
Closed

Use popcount intrinsincs in BitOperations#85944
kunalspathak wants to merge 3 commits into
dotnet:mainfrom
kunalspathak:popcnt

Conversation

@kunalspathak

Copy link
Copy Markdown
Contributor

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 8, 2023
@ghost

ghost commented May 8, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

I was trying to add this in #85842, but thought this should be evaluated separately on how much TP impact it makes.

Author:kunalspathak
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

I do see we replace the 2 steps comparison with popcnt.

image

The diffs are still off, but I don't think it will regress the execution time.
cc: @tannergooding

image

here is the analysis for minopts benchmarks.run windows/x64:

Base: 533135268, Diff: 533175333, +0.0075%
?BuildDefs@LinearScan@@AEAAXPEAUGenTree@@H_K@Z : 794531 : NA : 32.83% : +0.1490%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 304820 : +2.51% : 12.60% : +0.0572%
?BuildCall@LinearScan@@AEAAHPEAUGenTreeCall@@@Z : 123098 : +6.53% : 5.09% : +0.0231%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 3591 : +1.02% : 0.15% : +0.0007%
?genFnProlog@CodeGen@@IEAAXXZ : -2498 : -0.28% : 0.10% : -0.0005%
?BuildBlockStore@LinearScan@@AEAAHPEAUGenTreeBlk@@@Z : -3270 : -3.42% : 0.14% : -0.0006%
?PostOrderVisit@ForwardSubVisitor@@QEAA?AW4fgWalkResult@Compiler@@PEAPEAUGenTree@@PEAU4@@Z : -6030 : -9.86% : 0.25% : -0.0011%
?lvaAssignVirtualFrameOffsetsToLocals@Compiler@@QEAAXXZ : -23002 : -3.14% : 0.95% : -0.0043%
?genFinalizeFrame@CodeGen@@IEAAXXZ : -24372 : -32.17% : 1.01% : -0.0046%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : -168819 : -0.64% : 6.98% : -0.0317%
?BuildDef@LinearScan@@AEAAPEAVRefPosition@@PEAUGenTree@@_KH@Z : -460405 : -7.16% : 19.02% : -0.0864%
?BuildDefsWithKills@LinearScan@@AEAAXPEAUGenTree@@H_K1@Z : -499582 : -88.58% : 20.64% : -0.0937%

@kunalspathak
kunalspathak marked this pull request as ready for review May 9, 2023 05:06
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

#elif HOST_ARM64
return _CountOneBits(value);
#else
return __popcnt(value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't safe. It will always emit popcnt which requires SSE4.2

We'd need a cached CPUID check and a branch to use it, falling back to the bit twiddling logic if unsupported.

inline bool genExactlyOneBit(T value)
{
return ((value != 0) && genMaxOneBit(value));
return BitOperations::PopCount(value) == 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Given we need a branch to test for popcnt support, I expect the old logic may actually be faster as it generates:

 test edi, edi
jz false
lea eax, [rdi - 1]
test edi, eax
sete al
ret
false:
xor eax, eax
ret

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So existing logic has 2 branches vs. just 1 branch with popcount(). Do you think existing logic would still be faster?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing logic is just 1 branch. setcc is considered branchless and is specially handled by the CPU. The lea eax, [rdi - 1], test edi, eax, sete al is 3 cycles which is the same as for popcnt.

So it'd likely balance out, but with there now being 2 branches for anyone with "very old" hardware. I don't have a particular preference for which we do given that popcnt has been around for 15 years now, so we're unlikely to have people without it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it'd likely balance out

In that case, I don't think I should go ahead of using popcnt in genExactlyOneBit(), and just replace the existing bit twiddling logic with popcnt on supported hardware. That way the future BitOperations::PopCount() consumer will be optimized from popcnt. Thoughts?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That seems reasonable to me.

@BruceForstall

Copy link
Copy Markdown
Contributor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@kunalspathak Do you want to convert this to "Draft" so it doesn't get considered stale?

There is still work to be done to detect if popcnt is supported or not and if yes, then use it. I will mark it for "Ready" once that is done.

@JulieLeeMSFTJulieLeeMSFT added this to the 9.0.0 milestone Aug 7, 2023
@JulieLeeMSFT
JulieLeeMSFT marked this pull request as draft August 14, 2023 16:33
@ghostghost closed this Sep 13, 2023
@ghost

Copy link
Copy Markdown

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@ghostghost locked as resolved and limited conversation to collaborators Oct 13, 2023
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@tannergooding@JulieLeeMSFT