Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef - #90412

Merged
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers
Aug 16, 2023
Merged

Remove AggressiveOpt from CastHelpers.LdelemaRef/StelemRef#90412
EgorBo merged 10 commits into
dotnet:mainfrom
EgorBo:remove-aggressiveopt-casthelpers

Conversation

@EgorBo

@EgorBoEgorBo commented Aug 11, 2023

Copy link
Copy Markdown
Member

These two methods show up even for an empty app in JitDisasmSummary. The idea that we can treat them as normal managed methods and convert to direct calls once they reach Tier1 naturally (or FullOpts with TC=0).

Almost the final step (except the ETW stuff) to make empty apps jitting-free 🙂 (well, except the Main)

@ghostghost added the area-VM-coreclr label Aug 11, 2023
@ghostghost assigned EgorBoAug 11, 2023
@stephentoub

stephentoub commented Aug 11, 2023

Copy link
Copy Markdown
Member

And #90416 clears up the ETW stuff, if it's a valid fix (but it seems too easy 😄 ... we'll see what CI and others have to say)

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
@jkotas
jkotas requested review from VSadov and kouvelAugust 12, 2023 13:29
@EgorBo
EgorBo marked this pull request as ready for review August 12, 2023 17:22
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/tieredcompilation.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
Comment threadsrc/coreclr/vm/jitinterface.cpp
@EgorBo

EgorBo commented Aug 15, 2023

Copy link
Copy Markdown
MemberAuthor

This seems to improve benchmarks, e.g.:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){// stelem.ref_array[1]=_obj;_array[2]=_obj;}[Benchmark]publicboolTest2(){// CastHelpers.IsInstanceOf*return_objisIDisposable
or ICloneable
or Program;}
MethodJobToolchainMean
Test1Job-PCYDPA\runtime-main\corerun.exe3.531 ns
Test1Job-ZCRCGK\runtime\corerun.exe2.815 ns
Test2Job-PCYDPA\runtime-main\corerun.exe3.744 ns
Test2Job-ZCRCGK\runtime\corerun.exe3.525 ns

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same: https://www.diffchecker.com/RT66eY3O/ (stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Test2 is slightly faster because we now optimize indirect calls to direct ones: https://www.diffchecker.com/g5yrentL/ (the last call wasn't optimized to a direct one because it never reached the tier1 in this benchmark - it was never taken).

Hopefully, the difference is even better on arm64 (don't have any device around currently to test)

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr pgo, runtime-coreclr libraries-pgo

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@kouvel this is ready for the 2nd round 🙂 all PGO pipelines (with TC=1) look green. Benchmarks: #90412 (comment)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

I don't know why Test1 is faster but it seems like a stable improvement between runs, the codegen is the same

Possibly because AggressiveOptimization was removed and now it's using tier 1 code. There are code gen differences before and after this change in the helper method.

(stlem is a direct call as expected) - perhaps the method itself got allocated in a different part of the execution heap since it's no longer Aggressive opt? Or maybe the previous direct target wasn't entirely correct and pointed to a stub directly?

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

@kouvelkouvel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a comment/thought, LGTM though, thanks!

@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

Given that it was doing a direct call previously also, I wonder if there's a possibility that it could do a direct call to any tier including tier 0 (which would probably be r2r'd code, which may be ok perhaps) or an instrumented tier. Wonder if this should also be specialized for the other tiers to have it do an indirect call instead.

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine? Going to merge this one into 9.0 now, will see soon perf impact on microbenchmarks (my local runs show improvemnets)

The outerloop failure is #90593

@EgorBo
EgorBo merged commit 42862cc into dotnet:mainAug 16, 2023
@EgorBo
EgorBo deleted the remove-aggressiveopt-casthelpers branch August 16, 2023 00:55
@EgorBo

EgorBo commented Aug 16, 2023

Copy link
Copy Markdown
MemberAuthor

With this change, empty console app (without PublishReadyToRun) on Linux jits just one method now! Main itself:
image

(this PR removed these two cast helpers from this list + several PRs recently got rid of Vector<> paths in some BCL functions and made them prejittable + this is Linux where ETW doesn't trigger 🙂)

@kouvel

kouvel commented Aug 16, 2023

Copy link
Copy Markdown
Contributor

Like you mentioned, it had AggressiveOptimization attribute on it so presumably it was fine?

Yea I meant after this change, there would be R2R'ed code after the first call (AggressiveOptimization prevents R2R code to be generated), an intermediate code version with profiling instrumentation with TieredPGO, and finally the tier 1 code. The precode target I imagine would point to those things along the way and returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases. It is probably very rare though. Solving it would probably require taking the lock every time to check the tier.

Also similarly for the other cast helpers that don't have AggressiveOptimization.

@kouvel

Copy link
Copy Markdown
Contributor

Solving it would probably require taking the lock every time to check the tier.

Although there may be a shortcut to avoid the lock in startup cases where the helper method hasn't yet been called (maybe the precode target is not set to a native entry point).

@jkotas

Copy link
Copy Markdown
Member

returning the precode's target in the fallback path may mean that there is a possibility for there to be direct calls to non-final tiers in some cases

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

@kouvel

Copy link
Copy Markdown
Contributor

I do not think there is any code that would short-circuit the precode. If the JIT gets a pointer to precode, it is going to call it.

Ah ok, I misread the code, the fallback path would always be an indirect call and that would be fine. Thanks!

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Arm64 codegen diff:

object[]_array=newobject[1000];object_obj="test";[Benchmark]publicvoidTest1(){_array[1]=_obj;}[Benchmark]publicboolTest2(){return_objisIDisposable;}

https://www.diffchecker.com/WhfRiUit/

@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@EgorBo@stephentoub@kouvel@jkotas@VSadov