Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Force 2 iterations of OptRepeat - #100473

Closed
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease
Closed

Force 2 iterations of OptRepeat#100473
BruceForstall wants to merge 1 commit into
dotnet:mainfrom
BruceForstall:FixOptRepeat_ForceRepeat2InRelease

Conversation

@BruceForstall

Copy link
Copy Markdown
Contributor

No description provided.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Mar 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo

EgorBo commented Apr 1, 2024

Copy link
Copy Markdown
Member

Ouch, quite expensive (@SingleAccretion predicted similar numbers)
It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities.
Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Ouch, quite expensive (@SingleAccretion predicted similar numbers) It's still nice to have JitOptRepeat in good shape and it can point us to missing opportunities. Maybe there are some hidden fruits to reduce the TP? E.g.

  • Don't run it for huge methods (e.g. non-linear complexity for number of basic blocks, etc)
  • Run on demand? E.g. if RBO removes a branch and invalidates dominators or Assertprop propagates a constant and potentionally enables more opportunities for RBO?
  • Don't run some opts in the 2nd iter like assertprop and rangecheck if they don't contribute to the final diffs a lot

But of course that is quite non-trivial to guess whether you might benefit from the 2nd iter or not. Maybe if we study these diffs we'll figure out what we benefit the most from in most cases?

Lots of good ideas here. I agree that we wouldn't want to run it unconditionally. But running on big methods might be the most beneficial. And maybe we'd be willing to take the 20% TP cost for NAOT compiles.

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

cc @dotnet/jit-contrib

Example of enabling JitOptRepeat always, with 2 iterations. Code size improvement from about 3 to 4.5MB. TP cost of about 20%.

@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from bea3f5f to 19c5d79CompareApril 1, 2024 22:39
@BruceForstallBruceForstall changed the title Force OptRepeat, 2 iterations. Enable in Release.Force 2 iterations of OptRepeatApr 2, 2024
@BruceForstall
BruceForstallforce-pushed the FixOptRepeat_ForceRepeat2InRelease branch from 19c5d79 to 2c9dbdaCompareApril 2, 2024 02:04
@BruceForstallBruceForstall added the NO-MERGE The PR is not ready for merge yet (see discussion for detailed reasons) label Apr 2, 2024
@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Diffs

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

image
image
image

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here is a run of asmdiffs on win-x64 between JitOptRepeat 2 and 3 iterations, showing there is some benefit to be had to run more than twice.

Diffs are based on 2,384,245 contexts (931,974 MinOpts, 1,452,271 FullOpts).

MISSED contexts: base: 7,980 (0.33%), diff: 8,122 (0.34%)

Base JIT options: JitOptRepeat#*

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Overall (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,826-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch35,113,082-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch12,444,446-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch398,753,006-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,921,842-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,240,935-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch274,740,189-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,902,387-44,863+160.81%
realworld.run.windows.x64.checked.mch11,060,751-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,641-1,175-0.83%
FullOpts (-149,570 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,524,465-2,200-1.16%
benchmarks.run_pgo.windows.x64.checked.mch20,929,241-8,390+36.54%
benchmarks.run_tiered.windows.x64.checked.mch3,279,500-1,131-1.25%
coreclr_tests.run.windows.x64.checked.mch118,412,975-53,627+50.20%
libraries.crossgen2.windows.x64.checked.mch44,920,652-7,029-0.59%
libraries.pmi.windows.x64.checked.mch61,127,434-6,528+125.54%
libraries_tests.run.windows.x64.Release.mch102,077,791-23,431+20.96%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,588,969-44,863+160.81%
realworld.run.windows.x64.checked.mch10,647,587-1,196-0.12%
smoke_tests.nativeaot.windows.x64.checked.mch3,565,548-1,175-0.83%

@EgorBo

Copy link
Copy Markdown
Member

Sounds like 2 iterations are enough? 🙂 It'd be nice to have a breakdown TP report like @jakobbotsch did here for example to know what is the most expensive thing

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 3 and 4 iterations. There's still diffs.

Diffs are based on 2,384,210 contexts (931,974 MinOpts, 1,452,236 FullOpts).

MISSED contexts: base: 8,122 (0.34%), diff: 8,157 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#3

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Overall (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,626-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch35,104,692-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch12,443,315-295-0.03%
coreclr_tests.run.windows.x64.checked.mch398,699,105-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,914,813-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,219,837-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch274,716,321+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,840,633+11,773+374.53%
realworld.run.windows.x64.checked.mch11,059,555-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,466-484-0.19%
FullOpts (-71,360 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,522,265-1,154+0.26%
benchmarks.run_pgo.windows.x64.checked.mch20,920,851-3,889+70.31%
benchmarks.run_tiered.windows.x64.checked.mch3,278,369-295-0.03%
coreclr_tests.run.windows.x64.checked.mch118,359,074-64,162+420.88%
libraries.crossgen2.windows.x64.checked.mch44,913,623-3,739-0.10%
libraries.pmi.windows.x64.checked.mch61,106,336-8,283+306.51%
libraries_tests.run.windows.x64.Release.mch102,053,923+913+41.18%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,527,215+11,773+374.53%
realworld.run.windows.x64.checked.mch10,646,391-2,040-0.08%
smoke_tests.nativeaot.windows.x64.checked.mch3,564,373-484-0.19%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

Here's a run of asmdiffs on win-x64 between JitOptRepeat 4 and 5 iterations.

Diffs are based on 2,384,205 contexts (931,974 MinOpts, 1,452,231 FullOpts).

MISSED contexts: base: 8,157 (0.34%), diff: 8,162 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#4

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#5

Overall (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,472-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch35,100,803-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch12,443,020-318-0.19%
coreclr_tests.run.windows.x64.checked.mch398,634,943-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,911,074+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,206,674+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch274,717,234+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,838,676+5,179+577.00%
realworld.run.windows.x64.checked.mch11,057,515+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,982-261-0.02%
FullOpts (-14,538 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,521,111-388-0.12%
benchmarks.run_pgo.windows.x64.checked.mch20,916,962-1,606+79.30%
benchmarks.run_tiered.windows.x64.checked.mch3,278,074-318-0.19%
coreclr_tests.run.windows.x64.checked.mch118,294,912-28,742+118.37%
libraries.crossgen2.windows.x64.checked.mch44,909,884+2,441-0.31%
libraries.pmi.windows.x64.checked.mch61,093,173+4,410+746.88%
libraries_tests.run.windows.x64.Release.mch102,054,836+4,295+46.76%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,525,258+5,179+577.00%
realworld.run.windows.x64.checked.mch10,644,351+452+0.04%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,889-261-0.02%

@BruceForstall

Copy link
Copy Markdown
ContributorAuthor

And just "for fun", here is asmdiffs on win-x64 between JitOptRepeat 8 and 9 iterations! Yes, there are still diffs. One question: are optimization changes somehow monotonic, or is it possible the optimizations are oscillating (e.g., one iteration creates a CSE, another removes it)?

Diffs are based on 2,384,201 contexts (931,974 MinOpts, 1,452,227 FullOpts).

MISSED contexts: 8,166 (0.34%)

Base JIT options: JitOptRepeat#*;JitOptRepeatCount#8

Diff JIT options: JitOptRepeat#*;JitOptRepeatCount#9

Overall (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,520,257-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch35,095,669-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch12,442,539-38+0.00%
coreclr_tests.run.windows.x64.checked.mch398,475,394-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,910,655+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,193,216+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch274,706,838+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch135,827,290-2,776+1920.18%
realworld.run.windows.x64.checked.mch11,057,374+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,813+82+0.07%
FullOpts (+4,080 bytes)
CollectionBase size (bytes)Diff size (bytes)PerfScore in Diffs
benchmarks.run.windows.x64.checked.mch8,519,896-64-0.02%
benchmarks.run_pgo.windows.x64.checked.mch20,911,828-227+87.21%
benchmarks.run_tiered.windows.x64.checked.mch3,277,593-38+0.00%
coreclr_tests.run.windows.x64.checked.mch118,135,363-1,406+2445.47%
libraries.crossgen2.windows.x64.checked.mch44,909,465+1,299-0.03%
libraries.pmi.windows.x64.checked.mch61,079,715+2,633+2193.97%
libraries_tests.run.windows.x64.Release.mch102,044,440+3,977+61.81%
libraries_tests_no_tiered_compilation.run.windows.x64.Release.mch125,513,872-2,776+1920.18%
realworld.run.windows.x64.checked.mch10,644,210+600+0.00%
smoke_tests.nativeaot.windows.x64.checked.mch3,563,720+82+0.07%

@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 5, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMINO-MERGEThe PR is not ready for merge yet (see discussion for detailed reasons)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@BruceForstall@EgorBo