Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); JIT: Visit switch successors in increasing likelihood order for RPO-based layout by amanasifkhalid · Pull Request #101935 · dotnet/runtime · GitHub
Skip to content

JIT: Visit switch successors in increasing likelihood order for RPO-based layout - #101935

Closed
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout
Closed

JIT: Visit switch successors in increasing likelihood order for RPO-based layout#101935
amanasifkhalid wants to merge 3 commits into
dotnet:mainfrom
amanasifkhalid:greedy-switch-layout

Conversation

@amanasifkhalid

Copy link
Copy Markdown
Contributor

Follow-up to #101473. Also cleans up the BasicBlock successor visitor API surface a bit by separating logic for visiting successors in increasing likelihood order into BasicBlock::VisitAllSuccsInLikelihoodOrder.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

Here's the diff summary on win-x64, since the RPO-based layout is disabled by default. Diffs aren't that big, and they aren't overwhelmingly in one direction...

Diffs are based on 2,534,676 contexts (987,849 MinOpts, 1,546,827 FullOpts).

MISSED contexts: 2,922 (0.12%)

Overall (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch39,855,926+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,710,150-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch32,332,197-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch12,340,101-23+0.03%
coreclr_tests.run.windows.x64.checked.mch402,921,293-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,252,213-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,617,634-575-0.11%
libraries_tests.run.windows.x64.Release.mch285,052,618-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch137,049,028-574-0.14%
realworld.run.windows.x64.checked.mch13,552,182-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,032,760-861-1.60%
FullOpts (-8,298 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
aspnet.run.windows.x64.checked.mch27,307,698+1,073+0.20%
benchmarks.run.windows.x64.checked.mch8,709,728-108-0.05%
benchmarks.run_pgo.windows.x64.checked.mch18,519,682-718+0.39%
benchmarks.run_tiered.windows.x64.checked.mch3,077,579-23+0.03%
coreclr_tests.run.windows.x64.checked.mch121,562,176-325+0.27%
libraries.crossgen2.windows.x64.checked.mch45,250,508-1,591-0.13%
libraries.pmi.windows.x64.checked.mch63,504,133-575-0.11%
libraries_tests.run.windows.x64.Release.mch107,104,520-2,971+0.79%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch126,736,473-574-0.14%
realworld.run.windows.x64.checked.mch13,146,461-1,625-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch5,031,713-861-1.60%

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib, @AndyAyersMS PTAL. If you don't think it's worth churning the RPO layout implementation while running an experiment in the perf lab (especially since the diffs don't look like particularly large wins), I'm happy to undo those changes and just keep the block visitor cleanup work.

@amanasifkhalid

Copy link
Copy Markdown
ContributorAuthor

@jakobbotsch this cleans up the BasicBlock visitor API surface a bit. I initially tried something like the initializer pattern you mentioned over in #101473, but the current usage of AllSuccessorEnumerator in fgRunDfs complicated this approach -- what do you think of just passing the useProfile template argument in fgRunDfs to AllSuccessorEnumerator? It introduces some code duplication to the latter's constructor, but it's otherwise a pretty localized change.

public:
// Constructs an enumerator of all `block`'s successors.
AllSuccessorEnumerator(Compiler* comp, BasicBlock* block, const bool useProfile = false);
AllSuccessorEnumerator(Compiler* const comp, BasicBlock* const block);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: marking parameters as const in declarations is not very beneficial (it is a property of the definition, not of the declaration)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy: https://clang.llvm.org/extra/clang-tidy/checks/readability/avoid-const-params-in-decls.html

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a number of places where we have these const. IMO we should enable this rule of clang-tidy:

That seems reasonable. I can open a PR for that.

Comment on lines 735 to +739
Compiler::SwitchUniqueSuccSet sd = comp->GetDescriptorForSwitch(this);
jitstd::sort(sd.nonDuplicates, (sd.nonDuplicates + sd.numDistinctSuccs),
[](FlowEdge* const lhs, FlowEdge* const rhs) {
return (lhs->getLikelihood() * lhs->getDupCount()) < (rhs->getLikelihood() * rhs->getDupCount());
});

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems odd for the visitor function to be mutating the descriptor.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would make more sense to me if this sorting step was done ahead of time as part of the block layout phase.

Are you thinking of something like a pass over the block list that sorts all the switch descriptors, before computing the DFS? If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs, or adding another loop to the layout algorithm that won't do anything most of the time.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Something like that, or if that is undesirable, then some map created lazily in the context of the block layout phase (i.e. when we visit the switch block). I guess the latter is not much different to just allocating new memory to sort these successors into.
But at that point it makes more sense to me if we change AllSuccessorEnumerator slightly: in reality it's really just a vector with small vector optimization for N=4. If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If so, considering the diffs don't look all that significant, I'm tempted to just visit switch successors in their unsorted order to avoid cluttering the visitor APIs

I would suggest this, unless you have some information that order of successors matters for switches.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you phrase fgRunDfs in terms of a callback that returns this small vector of successors, then the memory where we can do the sorting without mutating existing data structures will already be available.

I tried something like this just now, but this ends up touching quite a bit of code if we want to keep the callback naive to the block kind (which seems desirable). I was thinking we'd adjust AllSuccessorEnumerator's underlying array to hold FlowEdge pointers instead of BasicBlock pointers to the block's successors, so that the callback can easily sort the array by edge likelihoods -- otherwise, we'd need to get the edge of each successor block with fgGetPredForBlock, which seems needlessly expensive. This means adjusting BasicBlock::VisitAllSuccs to take a callback that operates on FlowEdge* instead of BasicBlock*, but this doesn't work well in VisitEHSuccs, as EHblkDsc points to EH successors using BasicBlock pointers instead of FlowEdge pointers. We don't create edges to those successors, so we can't switch that data structure over to using edges.

I would suggest this, unless you have some information that order of successors matters for switches.

I think it makes most sense to leave switch successors alone for now, but still abstract the layout-specific ordering code out of VisitAllSuccs. If we're only doing anything special for BBJ_COND blocks for now, do we want a callback-based approach in fgRunDfs (I presume this callback would just manually compare the likelihoods of conditional blocks' true/false targets) that can be easily expanded later? Or are we ok with having a separate visitor method, VisitAllSuccsInLikelihoodOrder, that handles the BBJ_COND case while setting up the array in AllSuccessorEnumerator?

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amanasifkhalid@jakobbotsch@AndyAyersMS