Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Misc LSRA throughput improvements - #85842

Merged
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval
May 10, 2023
Merged

Misc LSRA throughput improvements#85842
kunalspathak merged 15 commits into
dotnet:mainfrom
kunalspathak:clearAssignedInterval

Conversation

@kunalspathak

@kunalspathakkunalspathak commented May 5, 2023

Copy link
Copy Markdown
Contributor

While working on consecutive-registers, I realized few things that could help in the throughput:
1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG. Done separately in #86016.

  1. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.

3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check. Done separately as part of #85944.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 5, 2023
@ghost

ghost commented May 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

While working on consecutive-registers, I realized few things that could help in the throughput:

  1. We pass around RegisterType in various methods, but that parameter is only used for TARGET_ARM. So wrap the parameter in `ARM_ARG.
  2. We call updateAssignedInterval() frequently, but more than half of the time, we pass interval == nullptr which is essentially clearing the interval. Introduced clearAssignedInterval() for that purpose.
  3. Use BitOperations::PopCount() in a method that is used for IsSingleRegister() check.
Author:kunalspathak
Assignees:kunalspathak
Labels:

area-CodeGen-coreclr

Milestone:-

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

TP regressions is surprising. Probably need to compare the assembly of before vs. after to see which individual change might be causing it.

Base: 538795279, Diff: 549102807, +1.9131%
?newRefPosition@LinearScan@@AEAAPEAVRefPosition@@PEAVInterval@@IW4RefType@@PEAUGenTree@@_KI@Z : 3693050 : +30.40% : 22.83% : +0.6854%
?associateRefPosWithInterval@LinearScan@@AEAAXPEAVRefPosition@@@Z : 3035189 : +29.77% : 18.76% : +0.5633%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 18.54% : +0.5566%
?applySelection@RegisterSelection@LinearScan@@AEAA_NH_K@Z : 2059105 : NA : 12.73% : +0.3822%
??$select@$0A@@RegisterSelection@LinearScan@@QEAA_KPEAVInterval@@PEAVRefPosition@@@Z : 1195021 : +4.50% : 7.39% : +0.2218%
?addRefsForPhysRegMask@LinearScan@@AEAAX_KIW4RefType@@_N@Z : 230276 : +3.90% : 1.42% : +0.0427%
?buildInternalRegisterUses@LinearScan@@AEAAXXZ : 28726 : +8.16% : 0.18% : +0.0053%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 18.01% : -0.5408%

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

uint32_tBitOperations::PopCount(uint32_t value)
{
#if defined(_MSC_VER)
// Inspired by the Stanford Bit Twiddling Hacks by Sean Eron Anderson:
// http://graphics.stanford.edu/~seander/bithacks.html
constuint32_t c1 = 0x55555555u;
constuint32_t c2 = 0x33333333u;
constuint32_t c3 = 0x0F0F0F0Fu;
constuint32_t c4 = 0x01010101u;
value -= (value >> 1) & c1;
value = (value & c2) + ((value >> 2) & c2);
value = (((value + (value >> 4)) & c3) * c4) >> 24;
return value;
#else
int32_t result = __builtin_popcount(value);
returnstatic_cast<uint32_t>(result);
#endif
}

newRefposition:

image

buildInternalregisterusage

image

I will revert the popcount change.

@BruceForstall

Copy link
Copy Markdown
Contributor

I didn't realize that we do not use intrinsics for popcount

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

@jakobbotsch

Copy link
Copy Markdown
Member

I didn't realize that we do not use intrinsics for popcount and which is why we are seeing lot of regressions. This is yet another example of why cross compilation comparison for TP might not be always accurate. cc: @jakobbotsch@BruceForstall

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

image

Looking at the diffs for minopts benchmarks_run windows-x64 I see slight regression:

Base: 538795279, Diff: 538880143, +0.0158%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@@Z : 2998888 : NA : 50.71% : +0.5566%
?updateAssignedInterval@LinearScan@@AEAAXPEAVRegRecord@@PEAVInterval@@W4var_types@@@Z : -2913849 : -100.00% : 49.27% : -0.5408%

But looking at the code, we should still profitable, because we are eliminating a condition:

image

@kunalspathak
kunalspathak marked this pull request as ready for review May 7, 2023 04:53
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Would be worthwhile profiling current function versus popcount -- probably we should switch to popcount.

Fixed.

@runfoapprunfoappBot mentioned this pull request May 8, 2023
@tannergooding

tannergooding commented May 8, 2023

Copy link
Copy Markdown
Member

It's a good point, but also a place where we should actively ensure we have parity regardless of the compiler we are using. It seems unfortunate that we are regressing either MSVC produced code or Clang produced code.

GCC/Clang generate the exact code that was codified for MSVC, they just do it implicitly via the builtin (and only switch to emitting actual popcnt if the ISA switch is passed in)

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

if the ISA switch is passed in

you mean when building clrjit using clang, right?

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

Agree. btw, I do see that with VC++ popcnt removes lot of that code and generate the actual intrinsic which will be faster too. I will send a separate PR to see the effect of that alone rather than mixing up with some of the LSRA improvements I am doing here.

@jakobbotsch

Copy link
Copy Markdown
Member

This was likely one of the many cases where we "execute more instructions" but the code was actually faster in practice.

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

GCC/Clang generate the exact code that was codified for MSVC

hhm. https://godbolt.org/z/rEndsvhv6

@tannergooding

Copy link
Copy Markdown
Member

I'm confused, are you saying the MSVC multiplication code was faster than using a single popcnt instruction?

@jakobbotsch: No, rather GCC/Clang don't emit popcnt here because our target machine is -msse2. In order for GCC/Clang to emit popcnt the target machine must be at least -msse42. For pre sse4.2, they emit the same logic as the multiplication code.

In order for us to emit popcnt here, we'd need to do it "opportunistically" via a cached CPUID check (much as we do for atomic operations on Arm64):

if (supportsPopcnt)
{
return __popcnt(value);
}
else
{
// Bit twiddling logic
}

@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

Latest diffs

image

@BruceForstall

Copy link
Copy Markdown
Contributor

Odd it shows x64 as a slight regression.

@kunalspathak
kunalspathak merged commit af1de13 into dotnet:mainMay 10, 2023
@kunalspathak
kunalspathak deleted the clearAssignedInterval branch May 10, 2023 05:00
@ghostghost locked as resolved and limited conversation to collaborators Jun 9, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kunalspathak@BruceForstall@jakobbotsch@tannergooding