Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); JIT: Move backward jumps to before their successors after RPO-based layout by amanasifkhalid · Pull Request #102461 · dotnet/runtime · GitHub
Skip to content

JIT: Move backward jumps to before their successors after RPO-based layout - #102461

Merged
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks
May 21, 2024
Merged

JIT: Move backward jumps to before their successors after RPO-based layout#102461
amanasifkhalid merged 6 commits into
dotnet:mainfrom
amanasifkhalid:move-hot-blocks

Conversation

@amanasifkhalid

@amanasifkhalidamanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
Contributor

Part of #93020. In #102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top (comment). To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.

Since the new layout isn't enabled in CI yet, here are the diffs with the baseline using the new RPO layout, on Windows x64: gist. This change is a net size regression, though the PerfScore diffs look promising, and I see many instances of the inverted loop shape fixed. TP impact is <=0.03% over doing just the RPO-based layout.

cc @dotnet/jit-contrib

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 20, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

New codegen for the example linked in this comment:

; Assembly listing for method Program:Sum(int[]):int (FullOpts)
; Emitting BLENDED_CODE for X64 with AVX - Windows
; FullOpts code
; optimized code
; rsp based frame
; fully interruptible
; No PGO data
; Final local variable assignments
;
; V00 arg0 [V00,T02] ( 4, 7 ) ref -> rcx class-hnd single-def <int[]>
;* V01 loc0 [V01,T05] ( 0, 0 ) int -> zero-ref
; V02 loc1 [V02,T04] ( 4, 6 ) int -> rax
; V03 OutArgs [V03 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace"
; V04 cse0 [V04,T01] ( 3, 10 ) int -> r10 "CSE #04: aggressive"
; V05 cse1 [V05,T03] ( 2, 9 ) int -> rdx hoist "CSE #01: aggressive"
; V06 rat0 [V06,T00] ( 5, 17 ) long -> r8 "Widened IV V01"
;
; Lcl frame size = 40
G_M53154_IG01: ;; offset=0x0000
sub rsp, 40
;; size=4 bbWeight=1 PerfScore 0.25
G_M53154_IG02: ;; offset=0x0004
xor eax, eax
mov edx, dword ptr [rcx+0x08]
xor r8d, r8d
jmp SHORT G_M53154_IG04
;; size=10 bbWeight=1 PerfScore 4.50
G_M53154_IG03: ;; offset=0x000E
add eax, r10d
inc r8d
;; size=6 bbWeight=2 PerfScore 1.00
G_M53154_IG04: ;; offset=0x0014
cmp edx, r8d
jle SHORT G_M53154_IG06
;; size=5 bbWeight=8 PerfScore 10.00
G_M53154_IG05: ;; offset=0x0019
mov r10d, dword ptr [rcx+4*r8+0x10]
test r10d, r10d
jne SHORT G_M53154_IG03
;; size=10 bbWeight=4 PerfScore 13.00
G_M53154_IG06: ;; offset=0x0023
add rsp, 40
ret
;; size=5 bbWeight=1 PerfScore 1.25
; Total bytes of code 40, prolog size 4, PerfScore 30.00, instruction count 14, allocated bytes for code 40 (MethodHash=477a305d) for method Program:Sum(int[]):int (FullOpts)
; ============================================================

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

I added a condition to fgMoveBlocksToHottestSuccessors to ensure loop heads are moved to the top of the loop, and updated the diffs link above. Adding this condition worsened size/PerfScore regressions, though I imagine we want to take this on principle? Here's the new layout for the System.Collections.Tests.Perf_BitArray.BitArrayNot(Size: 512) benchmark (notice BB11 now precedes BB12):

---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BBnum BBid ref try hnd preds weight IBC [IL range] [jump] [EH region] [flags]
---------------------------------------------------------------------------------------------------------------------------------------------------------------------
BB01 [0000] 1 1 897152 [000..039)-> BB23(0.001),BB08(0),BB07(0),BB06(0),BB05(0),BB04(0),BB03(0),BB02(0),BB09(0.999)[def] (switch) i IBC
BB09 [0009] 1 BB01 1.00 896255 [071..0D4)-> BB13(0),BB10(1) ( cond ) i IBC nullcheck
BB10 [0033] 1 BB09 1.00 896255 [0D6..???)-> BB12(1) (always) IBC internal
BB11 [0018] 1 BB12 15.66 14051500 [0D6..0F7)-> BB12(1) (always) i IBC loophead bwd bwd-target
BB12 [0019] 2 BB10,BB11 16.65 14939683 [0F7..107)-> BB11(0.941),BB21(0.0595) ( cond ) i IBC bwd bwd-src
BB21 [0028] 4 BB12,BB13,BB16,BB20 1.00 896255 [15A..15E)-> BB20(0),BB23(1) ( cond ) i IBC bwd bwd-src
BB23 [0029] 3 BB01,BB08,BB21 1.00 897152 [15E..16E) (return) i IBC
BB13 [0021] 1 BB09 0 0 [109..11A)-> BB21(0.48),BB16(0.52) ( cond ) i IBC rare
BB16 [0025] 2 BB13,BB15 0 0 [13D..14D)-> BB15(0.9),BB21(0.1) ( cond ) i IBC rare bwd bwd-src
BB15 [0024] 1 BB16 0 0 [11C..13D)-> BB16(1) (always) i IBC rare loophead bwd bwd-target
BB20 [0027] 1 BB21 0 0 [14F..15A)-> BB21(1) (always) i IBC rare loophead idxlen bwd bwd-target
BB02 [0002] 1 BB01 0 0 [03B..042)-> BB03(1) (always) i IBC rare idxlen
BB03 [0003] 2 BB01,BB02 0 0 [042..049)-> BB04(1) (always) i IBC rare idxlen
BB04 [0004] 2 BB01,BB03 0 0 [049..050)-> BB05(1) (always) i IBC rare idxlen
BB05 [0005] 2 BB01,BB04 0 0 [050..057)-> BB06(1) (always) i IBC rare idxlen
BB06 [0006] 2 BB01,BB05 0 0 [057..05E)-> BB07(1) (always) i IBC rare idxlen
BB07 [0007] 2 BB01,BB06 0 0 [05E..065)-> BB08(1) (always) i IBC rare idxlen
BB08 [0008] 2 BB01,BB07 0 0 [065..071)-> BB23(1) (always) i IBC rare idxlen
BB24 [0038] 0 0 [???..???) (throw ) i rare keep internal
---------------------------------------------------------------------------------------------------------------------------------------------------------------------

Edit: Looking at the diffs again, I don't think this exception is a good idea. Reverted.

@JulieLeeMSFTJulieLeeMSFT added the Priority:2 Work that is important, but not critical for the release label May 20, 2024
@AndyAyersMS

Copy link
Copy Markdown
Member

It's probably worth considering only BBJ_ALWAYS and BBJ_COND here, which makes finding the preferred successor simpler.

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

@amanasifkhalid

amanasifkhalid commented May 20, 2024

Copy link
Copy Markdown
ContributorAuthor

In the phoenix version the similar bit of code keyed off of RPO back edges. If a back edge source was not a BBJ_COND, and the target of the back edge was not following its most likely pred, we'd move the back edge source block up. Another way of saying this is that we'd really like the last block in the loop body to be an exit block.

I see, I think I can simplify this PR quite a bit then. I'm guessing Phoenix wouldn't do this for BBJ_COND blocks with hot back-edges because those could be loop exits? It might be too narrow in scope, but maybe fgMoveBlocksToHottestSuccessors should only consider BBJ_ALWAYS back-edges to fix any mis-rotated loops?

(also, sorry to loop you in while you're OOF)

@amanasifkhalidamanasifkhalid changed the title JIT: Move blocks up to their hottest successors after RPO-based layoutJIT: Move backward jumps to before their successors after RPO-based layoutMay 20, 2024
@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I've simplified my approach to only consider backward, unconditional jumps, and this seems to solve all the mis-rotated loop examples we looked at in #102343 while keeping churn and TP cost pretty low. This implementation is pretty narrow in the flowgraph shapes it's trying to fix, but since this work was motivated by the misshapen loop regressions, I think we can keep this conservative until we find a need for other passes to tweak the RPO-based order.

@EgorBo do you have any time to look at this today? If not, no worries; I think @jakobbotsch will be online tomorrow. Thanks!

@EgorBo

Copy link
Copy Markdown
Member

LGTM as far as I can tell, I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

I presume the newer commits you've pushed should still be zero diffs since it's RPO is not enabled, right?

Yes, that's right. I re-ran SPMI locally and updated the gist. TP impact is now <= 0.03%.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I think a simple targeted fix like this is just fine, but of course it would be nice with something a little more general in the future.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Thank you for the reviews! No diffs.

I saw you commented on #9304, but I wasn't sure what the final result ended up being; does this PR help there as well?

It does fix the loop rotation issue there as well, though the final layout isn't any better than what we started with. I'll elaborate more over there.

amanasifkhalid added a commit that referenced this pull request May 22, 2024
Follow-up to #102461, and part of #9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see #9304 for one such example.
steveharter pushed a commit to steveharter/runtime that referenced this pull request May 28, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…ayout (dotnet#102461)
Part of dotnet#93020. In dotnet#102343, we noticed the RPO-based layout sometimes makes suboptimal decisions in terms of placing a block's hottest predecessor before it -- in particular, this affects loops that aren't entered at the top. To address this, after establishing a baseline RPO layout, fgMoveBackwardJumpsToSuccessors will try to move backward unconditional jumps to right behind their targets to create fallthrough, if the predecessor block is sufficiently hot.
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
Follow-up to dotnet#102461, and part of dotnet#9304. Compacting blocks after establishing an RPO-based layout, but before moving backward jumps to fall into their successors, can enable more opportunities for branch removal; see dotnet#9304 for one such example.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIPriority:2Work that is important, but not critical for the release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@amanasifkhalid@AndyAyersMS@EgorBo@jakobbotsch@JulieLeeMSFT