Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Improve Guid equality checks on 64-bit platforms by BruceForstall · Pull Request #35654 · dotnet/runtime · GitHub
Skip to content

Improve Guid equality checks on 64-bit platforms - #35654

Closed
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks
Closed

Improve Guid equality checks on 64-bit platforms#35654
BruceForstall wants to merge 1 commit into
dotnet:masterfrom
BruceForstall:ImproveGuidEqualityChecks

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.

This will be even better on arm64 which passes everything
in 64-bit registers, after #35622 is addressed in the JIT.

Perf results:

x64:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase2.322 ns0.0210 ns0.0187 ns2.324 ns2.294 ns2.351 ns1.000.00
EqualsSamediff1.547 ns0.0092 ns0.0071 ns1.547 ns1.535 ns1.559 ns0.670.01
EqualsOperatorbase2.890 ns0.3896 ns0.4169 ns3.030 ns1.722 ns3.074 ns1.000.00
EqualsOperatordiff1.346 ns0.0160 ns0.0150 ns1.346 ns1.331 ns1.380 ns0.490.12
NotEqualsOperatorbase1.738 ns0.0306 ns0.0255 ns1.730 ns1.712 ns1.805 ns1.000.00
NotEqualsOperatordiff1.401 ns0.0425 ns0.0355 ns1.389 ns1.360 ns1.476 ns0.810.02

x86:

MethodToolMeanErrorStdDevMedianMinMaxRatioRatioSD
EqualsSamebase3.164 ns0.0234 ns0.0208 ns3.159 ns3.136 ns3.203 ns1.000.00
EqualsSamediff3.079 ns0.0327 ns0.0306 ns3.074 ns3.041 ns3.146 ns0.970.01
EqualsOperatorbase2.736 ns0.0252 ns0.0236 ns2.726 ns2.710 ns2.783 ns1.000.00
EqualsOperatordiff2.613 ns0.0262 ns0.0245 ns2.600 ns2.589 ns2.662 ns0.950.01
NotEqualsOperatorbase2.708 ns0.0096 ns0.0080 ns2.705 ns2.699 ns2.723 ns1.000.00
NotEqualsOperatordiff2.573 ns0.0666 ns0.0591 ns2.552 ns2.526 ns2.709 ns0.950.02

Current code does four 32-bit comparisons. Instead,
do two 64-bit comparisons. On x86, the JIT-generated
code is slightly different, but equally fast.
This will be even better on arm64 which passes everything
in 64-bit registers, after dotnet#35622 is addressed in the JIT.
Perf results:
x64:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 2.322 ns | 0.0210 ns | 0.0187 ns | 2.324 ns | 2.294 ns | 2.351 ns | 1.00 | 0.00 |
| EqualsSame | diff | 1.547 ns | 0.0092 ns | 0.0071 ns | 1.547 ns | 1.535 ns | 1.559 ns | 0.67 | 0.01 |
| EqualsOperator | base | 2.890 ns | 0.3896 ns | 0.4169 ns | 3.030 ns | 1.722 ns | 3.074 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 1.346 ns | 0.0160 ns | 0.0150 ns | 1.346 ns | 1.331 ns | 1.380 ns | 0.49 | 0.12 |
| NotEqualsOperator | base | 1.738 ns | 0.0306 ns | 0.0255 ns | 1.730 ns | 1.712 ns | 1.805 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 1.401 ns | 0.0425 ns | 0.0355 ns | 1.389 ns | 1.360 ns | 1.476 ns | 0.81 | 0.02 |
x86:
| Method | Tool | Mean | Error | StdDev | Median | Min | Max | Ratio | RatioSD |
|------------------ |------|-----------:|----------:|----------:|-----------:|-----------:|-----------:|------:|--------:|
| EqualsSame | base | 3.164 ns | 0.0234 ns | 0.0208 ns | 3.159 ns | 3.136 ns | 3.203 ns | 1.00 | 0.00 |
| EqualsSame | diff | 3.079 ns | 0.0327 ns | 0.0306 ns | 3.074 ns | 3.041 ns | 3.146 ns | 0.97 | 0.01 |
| EqualsOperator | base | 2.736 ns | 0.0252 ns | 0.0236 ns | 2.726 ns | 2.710 ns | 2.783 ns | 1.00 | 0.00 |
| EqualsOperator | diff | 2.613 ns | 0.0262 ns | 0.0245 ns | 2.600 ns | 2.589 ns | 2.662 ns | 0.95 | 0.01 |
| NotEqualsOperator | base | 2.708 ns | 0.0096 ns | 0.0080 ns | 2.705 ns | 2.699 ns | 2.723 ns | 1.00 | 0.00 |
| NotEqualsOperator | diff | 2.573 ns | 0.0666 ns | 0.0591 ns | 2.552 ns | 2.526 ns | 2.709 ns | 0.95 | 0.02 |
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Additional benchmarks: dotnet/performance#1302

@EgorBo

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

@gfoidl

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

@EgorBo

Copy link
Copy Markdown
Member

Did you consider a vectorized approach?
Only shown for SSE2, needs some specialization for Arm, and I don't know if it's worth it...

Last time I tried it was slower,
I think in most cases guids are different and early out check is enough.

Unsafe.Add(ref g._a, 1) == Unsafe.Add(ref _a, 1) &&
Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this logic need to be repeated in multiple places rather than have everything forward to == for example?

I'd hope the JIT would properly inline the check in all the cases and the code would be equivalent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, although that's a separate issue.

Unsafe.Add(ref g._a, 2) == Unsafe.Add(ref _a, 2) &&
Unsafe.Add(ref g._a, 3) == Unsafe.Add(ref _a, 3);
return Unsafe.As<int, long>(ref g._a) == Unsafe.As<int, long>(ref _a) &&
Unsafe.As<byte, long>(ref g._d) == Unsafe.As<byte, long>(ref _d);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was under the impression that Int64-based operations would be slower than Int32-based operations on 32-bit, especially if the values weren't aligned (and that that's why @jkotas used Int32 here initially when switching these operations to be based on Unsafe). Is that not the case?

@jkotas

Copy link
Copy Markdown
Member

Is it safe for arm32 as well (accessing misaligned long)?

+1. It would be better to use Unsafe.ReadUnaligned here to make the code portable.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Did you consider a vectorized approach?

I didn't. I was specifically looking to improve arm64, but x64 as well (and be simple and cross-platform).

Is it safe for arm32 as well (accessing misaligned long)?

RyuJIT converts a 64-bit long access to two 32-bit int accesses on 32-bit platforms, so there is no difference in alignment for 32-bit. There could be a difference in alignment for 64-bit since presumably Guid could be 4-byte aligned.

fyi, here's the x86 code difference. It appears RyuJIT doesn't handle the Unsafe.Add calls well.

current x86 assembly
G_M51749_IG01:55pushebp 8BEC movebp,esp ;; bbWeight=1 PerfScore 1.25G_M51749_IG02: 8B4518 moveax, dword ptr [ebp+18H] 3B4508 cmpeax, dword ptr [ebp+08H]7532jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 3.00G_M51749_IG03: 8D4518 leaeax, bword ptr [ebp+18H] 8B4004 moveax, dword ptr [eax+4] 8D5508 leaedx, bword ptr [ebp+08H] 3B4204 cmpeax, dword ptr [edx+4]7524jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B4008 moveax, dword ptr [eax+8] 8D5508 leaedx, bword ptr [ebp+08H] 3B4208 cmpeax, dword ptr [edx+8]7516jne SHORT G_M51749_IG05 8D4518 leaeax, bword ptr [ebp+18H] 8B400C moveax, dword ptr [eax+12] 8D5508 leaedx, bword ptr [ebp+08H] 3B420C cmpeax, dword ptr [edx+12] 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 9.13G_M51749_IG04: 5D popebp C22000 ret32 ;; bbWeight=0.50 PerfScore 1.25G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: 5D popebp C22000 ret32
new x86 assembly
G_M51749_IG02: 8B442414 moveax, dword ptr [esp+14H] 8B542418 movedx, dword ptr [esp+18H]33442404xoreax, dword ptr [esp+04H]33542408xoredx, dword ptr [esp+08H] 0BC2 oreax,edx 751B jne SHORT G_M51749_IG05 ;; bbWeight=1 PerfScore 5.25G_M51749_IG03: 8B44241C moveax, dword ptr [esp+1CH] 8B542420 movedx, dword ptr [esp+20H] 3344240C xoreax, dword ptr [esp+0CH]33542410xoredx, dword ptr [esp+10H] 0BC2 oreax,edx 0F94C0 sete al 0FB6C0 movzxeax,al ;; bbWeight=0.50 PerfScore 2.75G_M51749_IG04: C22000 ret32 ;; bbWeight=0.50 PerfScore 1.00G_M51749_IG05: 33C0 xoreax,eax ;; bbWeight=0.50 PerfScore 0.13G_M51749_IG06: C22000 ret32

@stephentoub

Copy link
Copy Markdown
Member

@BruceForstall, are you still working on this? Thanks.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

@BruceForstall, are you still working on this? Thanks.

I haven't, and likely won't get back to it. I'll go ahead and close the PR. I'd be happy if someone picked it up. The "only thing" is to try Jan's suggestion about using Unsafe.ReadUnaligned, and verifying performance on all platforms still improves.

@jkotasjkotas mentioned this pull request Nov 13, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall restored the ImproveGuidEqualityChecks branch February 28, 2021 03:08
@BruceForstall
BruceForstall deleted the ImproveGuidEqualityChecks branch December 28, 2022 01:04
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@BruceForstall@EgorBo@gfoidl@jkotas@stephentoub@tannergooding@Dotnet-GitSync-Bot