JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Compute BB weights with higher precision, use tolerance in if-conversion - #88385

Merged
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance
Jul 5, 2023
Merged

JIT: Compute BB weights with higher precision, use tolerance in if-conversion#88385
jakobbotsch merged 3 commits into
dotnet:mainfrom
jakobbotsch:if-conversion-weight-tolerance

Conversation

@jakobbotsch

Copy link
Copy Markdown
Member

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.

…nversion
In dotnet#88376 (comment)
I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex
benchmark that looked like this:
```diff
G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx
+ jg SHORT G_M62987_IG15
+ ;; size=12 bbWeight=1.00 PerfScore 2.50
+G_M62987_IG14:
+ mov ebx, 1
+ ;; size=5 bbWeight=0.98 PerfScore 0.24
+G_M62987_IG15:
```
Investigating, it turns out to be caused by a check in if-conversion for
detecting loops, that is using a very sharp boundary. It turns out that
the integral block weights are exactly the same (in both the base and
diff), but the floating point calculation ended up with exactly 1 and
very close to 1 in the base/diff cases respectively.
So two changes to address it:
* Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
* Switch if-conversion to have a 5% tolerance so that losing just a
single count in the wrong place will not cause a decision change. Note
that the check here is used as a cheap "is inside loop" check.
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 4, 2023
@ghostghost assigned jakobbotschJul 4, 2023
@ghost

ghost commented Jul 4, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

In #88376 (comment) I noticed a spurious diff in the System.Tests.Perf_Int32.ToStringHex benchmark that looked like this:

 G_M62987_IG13:
and edi, esi
mov ebx, dword ptr [rbp+88H]
- mov ecx, 1
test ebx, ebx
- cmovle ebx, ecx+ jg SHORT G_M62987_IG15+ ;; size=12 bbWeight=1.00 PerfScore 2.50+G_M62987_IG14:+ mov ebx, 1+ ;; size=5 bbWeight=0.98 PerfScore 0.24+G_M62987_IG15:

Investigating, it turns out to be caused by a check in if-conversion for detecting loops, that is using a very sharp boundary. It turns out that the integral block weights are exactly the same (in both the base and diff), but the floating point calculation ended up with exactly 1 and very close to 1 in the base/diff cases respectively.

So two changes to address it:

  • Switch getBBWeight to multiply by BB_UNITY_WEIGHT after dividing
  • Switch if-conversion to have a 5% tolerance so that losing just a single count in the wrong place will not cause a decision change. Note that the check here is used as a cheap "is inside loop" check.
Author:jakobbotsch
Assignees:jakobbotsch
Labels:

area-CodeGen-coreclr

Milestone:-

Comment threadsrc/coreclr/jit/block.cpp Outdated
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to stylize these comparisons, so maybe add a new fgProfileWeight... helper here?

In particular these sorts of checks are often parts of profitability heuristics which we might want to vary under stress and/or expose to some machine learning mechanism.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think playing around with this threshold makes much sense given its use as an "is inside loop" check. The precise check happens below when this check is false.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not really checking if the block is likely to be inside a loop, it is checking if the block is likely to be inside a frequently executed loop.

If we assume profile data is representative, then perhaps this is fine; if we think it might not be representative, then perhaps we should be checking bbNatLoopNum and if set, either bypass the conversion, or go on to compare the weight of this block to the weight of the loop entry (so either "in a loop" or "frequently executed wrt the loop"). I wonder if this would subsume the need for an optReachable check.

Do we know if the cases in #82106 did not get handled by our loop recognition? I don't recall why we didn't check this initially -- do we no longer trust the loop table? Seems like it might still be good enough.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression in #82106 was because we if-converted inside an unnatural loop that was not recognized by our loop recognition. The loop is this code:

do
{
// Make sure we don't go out of bounds
Debug.Assert(offset+ch1ch2Distance+Vector256<ushort>.Count<=searchSpaceLength);
Vector256<ushort>cmpCh2=Vector256.Equals(ch2,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)(offset+ch1ch2Distance)));
Vector256<ushort>cmpCh1=Vector256.Equals(ch1,Vector256.LoadUnsafe(refushortSearchSpace,(nuint)offset));
Vector256<byte>cmpAnd=(cmpCh1&cmpCh2).AsByte();
// Early out: cmpAnd is all zeros
if(cmpAnd!=Vector256<byte>.Zero)
{
gotoCANDIDATE_FOUND;
}
LOOP_FOOTER:
offset+=Vector256<ushort>.Count;
if(offset==searchSpaceMinusValueTailLength)
return-1;
// Overlap with the current chunk for trailing elements
if(offset>searchSpaceMinusValueTailLengthAndVector)
offset=searchSpaceMinusValueTailLengthAndVector;
continue;
CANDIDATE_FOUND:
uintmask=cmpAnd.ExtractMostSignificantBits();
do
{
intbitPos=BitOperations.TrailingZeroCount(mask);
// div by 2 (shr) because we work with 2-byte chars
nintcharPos=(nint)((uint)bitPos/2);
if(valueLength==2||// we already matched two chars
SequenceEqual(
refUnsafe.As<char,byte>(refUnsafe.Add(refsearchSpace,offset+charPos)),
refUnsafe.As<char,byte>(refvalue),(nuint)(uint)valueLength*2))
{
return(int)(offset+charPos);
}
// Clear two the lowest set bits
if(Bmi1.IsSupported)
mask=Bmi1.ResetLowestSetBit(Bmi1.ResetLowestSetBit(mask));
else
mask&=~(uint)(0b11<<bitPos);
}while(mask!=0);
gotoLOOP_FOOTER;
}while(true);

That's why we ended up with optReachable here.

There is an alternative to simply get rid of this block weight check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

Co-authored-by: Andy Ayers <andya@microsoft.com>
// Detect via the block weight as that will be high when inside a loop.
if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT)

if (m_startBlock->getBBWeight(m_comp) > BB_UNITY_WEIGHT * 1.05)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok... let's leave this as is for now.

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

@BruceForstall

Copy link
Copy Markdown
Contributor

@BruceForstall do we have a work item for loops we ought to recognize but don't? I see #43713 so perhaps that's good enough for now.

That's the issue I'd reference

@jakobbotsch
jakobbotsch merged commit 3f19af8 into dotnet:mainJul 5, 2023
@jakobbotsch
jakobbotsch deleted the if-conversion-weight-tolerance branch July 5, 2023 18:37
@ghostghost locked as resolved and limited conversation to collaborators Aug 4, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakobbotsch@BruceForstall@AndyAyersMS