Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Optimize System.HexConverter.IsHexChar on 64 bits - #52470

Merged
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt
May 11, 2021
Merged

Optimize System.HexConverter.IsHexChar on 64 bits#52470
GrabYourPitchforks merged 3 commits into
dotnet:mainfrom
Sergio0694:is-hex-digit-opt

Conversation

@Sergio0694

@Sergio0694Sergio0694 commented May 7, 2021

Copy link
Copy Markdown
Contributor

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Add a branchless fast path on 64 bit systems that doesn't do memory accesses either
@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

I couldn't figure out the best area label to add to this PR. If you have write-permissions please help me learn by adding exactly one area label.

public static bool IsHexChar(int c)
public static unsafe bool IsHexChar(int c)
{
if (sizeof(IntPtr) == 8)

@EgorBoEgorBoMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since corelib is always arch specific I guess it's better (e.g. for JIT) to use #if TARGET_64BIT here

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is in the Common folder and is pulled in by a bunch of different libs, some of which only build AnyCPU. If we wanted an ifdef for the target platform, we'd also need to ifdef CORELIB. (See ValueStringBuilder.cs for some examples.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@GrabYourPitchforks ah, good point!

shift = unchecked((int)((35465847073801215UL >> i) & 1)),
and = shift & mask;
byte result = unchecked((byte)and);
bool valid = *(bool*)&result;

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We considered similar tricks to this a while back and those tricks were rejected. See dotnet/coreclr#16138 for one example.
The tl;dr was "somebody's probably branching on the result of this method, so it's better if the final statement is an actual comparison that can be folded into the caller's branch condition."

The implication would be that these lines become return (and & 1) != 0;, or return (and & 1) != 0 ? true : false; if we're trying to work around #4207.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes sense, I was thinking the last return expression could be changed to that due to callers branching 🙂
Question though: given the input is guaranteed to be either 1 or 0 here due to the previous logic, we can just test and != 0 here, right? As in, we can skip that & 1 since that wouldn't affect the actual semantics in this case, no?

So at the end of the day:

returnand!=0?true:false;

Should work?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yup, that should work!

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: should read ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused, aren't we checking that c is also >= 0 here?
From your original snippet on Discord:

ulongmask=i-64;// c is negative IFF '0' <= c < '0' + 64; else c is non-negative

Am I mixing things up here? 🤔

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If c is within [0, '0'), bit 31 will be set, but bit 63 (the bit we care about for masking purposes) will not be set.
Bit 63 only gets set if c is within ['0', '0' + 64).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh right, yeah. Fixed, thanks! 😄

Comment threadsrc/libraries/Common/src/System/HexConverter.cs Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove unsafe keyword; no longer needed.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment below, done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5.

@GrabYourPitchforksGrabYourPitchforksMay 7, 2021

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prefer IntPtr.Size rather than sizeof(IntPtr) here. illink can see this as a proper method call to IntPtr.get_Size and will perform substitution. See, e.g., https://github.com/dotnet/runtime/blob/74cafe71de68fac46cd4956da98387bce9e1043c/src/libraries/System.Private.CoreLib/src/ILLink/ILLink.Substitutions.64bit.xml.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, TIL. Done in 78bc559532402d2d4dd1a2e9a85b9e9efd4240a5 🙂

@GrabYourPitchforksGrabYourPitchforks left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Appreciate the perf numbers in the issue as well.

@ghost

ghost commented May 7, 2021

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
See info in area-owners.md if you want to be subscribed.

Issue Details

Overview

This PR adds a fast path to System.HexConverter.IsHexChar(int) on 64 bit systems, which has:

  • No branches (down from 2 conditional + 1 unconditional)
  • Smaller codegen (66 bytes ---> 35 bytes)
  • No memory accesses (so the speed is no longer affected by cache state)

The change is a specialized version of what I used in BitHelper.HasLookupFlag in the Microsoft.Toolkit.HighPerformance package, and just uses bit trickery to make the code branchless and read the lookup value from a constant value and not from memory.

Codegen diff

Before (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar_OG(int):boolG_M3768_IG01:subrsp,40 ;; bbWeight=1 PerfScore 0.25G_M3768_IG02:cmpecx,256jge SHORT G_M3768_IG04 ;; bbWeight=1 PerfScore 1.25G_M3768_IG03:cmpecx,256jae SHORT G_M3768_IG07movsxdrax,ecxmovrdx,0xD1FFAB1Emovzxrax, byte ptr [rax+rdx]jmp SHORT G_M3768_IG05 ;; bbWeight=0.50 PerfScore 2.88G_M3768_IG04:moveax,255 ;; bbWeight=0.50 PerfScore 0.12G_M3768_IG05:cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 1.50G_M3768_IG06:addrsp,40ret ;; bbWeight=1 PerfScore 1.25G_M3768_IG07:call CORINFO_HELP_RNGCHKFAILint3 ;; bbWeight=0 PerfScore 0.00; Total bytes of code: 66
After (click to expand):
; Method IsHexCharFast.HexConverter2:IsHexChar(int):boolG_M2063_IG01: ;; bbWeight=1 PerfScore 0.00G_M2063_IG02:addecx,-48learax,[rcx-64]movrdx,0xD1FFAB1Eshlrdx,clandrax,rdxjl SHORT G_M2063_IG05 ;; bbWeight=1 PerfScore 4.25G_M2063_IG03:xoreax,eax ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG04:ret ;; bbWeight=0.50 PerfScore 0.50G_M2063_IG05:moveax,1 ;; bbWeight=0.50 PerfScore 0.12G_M2063_IG06:ret ;; bbWeight=0.50 PerfScore 0.50; Total bytes of code: 34

Additionally, the new version can also be JITted to just a constant, if the input is a constant.
That is, if you call IsHexChar with an input constant like 'A', you get this JIT diff:

Before (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_OG():boolG_M36347_IG01: ;; bbWeight=1 PerfScore 0.00G_M36347_IG02:movrax,0xD1FFAB1Emovzxrax, byte ptr [rax]cmpeax,255 setne almovzxrax,al ;; bbWeight=1 PerfScore 3.75G_M36347_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 25
After (click to expand):
; Method IsHexCharFast.HexConverter2:Check_Constant_NEW():boolG_M50831_IG01: ;; bbWeight=1 PerfScore 0.00G_M50831_IG02:moveax,1 ;; bbWeight=1 PerfScore 0.25G_M50831_IG03:ret ;; bbWeight=1 PerfScore 1.00; Total bytes of code: 6

Benchmark

I've put together a small test benchmark which you can find here.

MethodMeanErrorStdDevRatioCode Size
IsHexChar_OG_Random9.230 us0.0508 us0.0475 us1.0072 B
IsHexChar_NEW_Random4.406 us0.0331 us0.0309 us0.4862 B
IsHexChar_OG_AlwaysTrue2.934 us0.0183 us0.0172 us1.0072 B
IsHexChar_NEW_AlwaysTrue2.947 us0.0199 us0.0186 us1.0062 B

The new version is about 2x faster over random data in this test. When the input is always valid and the branch predictor can be more effective in the original implementation, this test shows the new version is still on par, but a couple notes:

  • This benchmark is not representative of real world data since the method is just constantly being invoked in a loop, so the original version will just do cache hits every single time, which gives it an advantage here.
  • Even not taking this into consideration, the new version is just generally not dependent on input data and always produces consistent performance in all cases (which I would argue is better even with performance being the same in this test)

JIT diff

Currently work in progress, I'm not having luck with the runtime tooling today... 😄
Opened the PR in the meantime to have the CI run on it at least.

Author:Sergio0694
Assignees:-
Labels:

area-System.Runtime

Milestone:-

@GrabYourPitchforks

Copy link
Copy Markdown
Member

@EgorBo@tannergooding any other thoughts on this? Thinking of merging it in Tuesday, which gives a little while longer for comments.

@GrabYourPitchforks
GrabYourPitchforks merged commit 0ee10e9 into dotnet:mainMay 11, 2021
@Sergio0694
Sergio0694 deleted the is-hex-digit-opt branch May 11, 2021 19:16
@danmoseley

Copy link
Copy Markdown
Contributor

Nice contribution @Sergio0694 thank you.

@karelzkarelz added this to the 6.0.0 milestone May 20, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jun 19, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Sergio0694@GrabYourPitchforks@danmoseley@EgorBo@karelz