Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add support for (experimental) profile-based trimming - #108049

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim
Nov 11, 2024
Merged

Add support for (experimental) profile-based trimming#108049
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
MichalStrehovsky:instrtrim

Conversation

@MichalStrehovsky

Copy link
Copy Markdown
Member

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.

The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.

The idea is simple:

  • We build a special version of the program that keeps track of what methods executed.
  • At process termination, this data is written out to a file.
  • We then recompile the program, passing in the method list as one of the inputs.
  • When compiling a method that was in the list, compile as usual.
  • When compiling a method that wasn’t in the list, replace it with a failfast.
  • We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.

The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with -p:Deterministic=true.

Usage

Build the app with -p:_InstrumentReachability=true. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.

Then build the app again with -p:_ReachabilityInstrumentationFile=reach.mprof. Output of this build will be a profile-based version.

Just to give an idea of how small things will be:

  • dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
  • Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB

The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:

typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();classMyAttribute:Attribute{publicMyAttribute(Typet)=>TheType=t;publicTypeTheType;}classGen<T>{}[My(typeof(Gen<>))]classHolder{}

(We need to call IsConstructedGenericType on a type that got its MethodTable optimized away.)

It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.

Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).

Implementation

There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.

Generating profile data

ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).

ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.

Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.

Saving profile data

Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.

Consuming profile data

This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.

Cc @dotnet/ilc-contrib

This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
Member

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

@kekekeks

Copy link
Copy Markdown

Will profile be only available with NAOT builds? i. e. would one still have to make sure that the trimmed app still runs with NAOT first to collect the profile?

I was thinking about generating such profile for non-corlib stuff using CoreCLR profiling APIs while forcing various codepaths to think that the code runs with NAOT by IL-patching IsDynamicCodeSupported and RuntimeInformation.FrameworkDescription.

@kekekeks

Copy link
Copy Markdown

I guess two-stage profiling would be quite useful. A CoreCLR-based one to make the app to run with NAOT in the first place (without having to spend several hours on adjusting trimming configuration by trial and error) and then the precise one from instrumented NativeAOT build to produce a smaller binary.

@GerardSmit

GerardSmit commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

  1. Create a web application
  2. Create a new controller
  3. Visit the controller, generate the reach.mprof-file and commit it to git
  4. Create a new controller
  5. Only visit the new controller and let the profile data merge into the reach.mprof-file

That both controllers still get included (including all the underlying actions, which could run more code like DB access)?
Or do you need to recreate the reach.mprof-file every time (visit both controllers before publishing)

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

Do you plan to use MIBC for this? (so not only you can extract the reachability info, but also improve performance with PGO)

This only logs what methods run, no basic blocks. I went for 20% of effort and 80% of effect. If we ever have proper profile collection, half of this could probably be deleted, we wouldn't need the instrumentation and corelib part of this.

You can run the app multiple times, the profile data will get merged automatically.

Does this mean that:

You can run it multiple times, you cannot recompile it. The file format uses tokens and MVIDs. If you change stuff, they get shuffled and update is rejected.


To be very clear, the only purpose of this is to:

  1. Find out the best possible scenario when it comes to size (how much more one could save in the very ideal and unrealistic case)
  2. Find out if there are any things that could be factored differently so that regular trimming can get rid of them.

Do not ever ship anything compiled like this, it explodes randomly (e.g. if you never profiled contended lock situation and a lock in your app becomes contended, the app will just crash). I even did the extra effort to start the MSBuild properties that activate this with underscores to deter anyone from checking in code that has this.

@EgorBo

EgorBo commented Sep 20, 2024

Copy link
Copy Markdown
Member

This only logs what methods run, no basic blocks.

MIBC is expected to collect data about basic blocks, it's just that by default it may skip some blocks/methods to make it lightweight. That can be changed via

DOTNET_JitEdgeProfiling=0
DOTNET_JitMinimalJitProfiling=0

I went for 20% of effort and 80% of effect.

Understandable

@am11am11 added the Hackathon Issues picked for Hackathon label Sep 21, 2024

@agockeagocke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did a group code review and I asked most questions there, so LGTM thanks

Comment on lines +240 to +246
public override TypeSystemContext Context
{
get
{
return _context;
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI you could use an auto-property to make these shorter if you want, e.g.

Suggested change
publicoverrideTypeSystemContextContext
{
get
{
return_context;
}
}
publicoverrideTypeSystemContextContext{get;}

@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/ba-g build took too long and timed out in an unrelated leg

@MichalStrehovsky
MichalStrehovsky merged commit 2c29c1d into dotnet:mainNov 11, 2024
@MichalStrehovsky
MichalStrehovsky deleted the instrtrim branch November 11, 2024 14:50
mikelle-rogers pushed a commit to mikelle-rogers/runtime that referenced this pull request Dec 10, 2024
* Add support for (experimental) profile-based trimming
This was my hackathon project of this year. Since it’s in a pretty leaf-y location within the product, I think it would be fine to check it in under unsupported switches for experimentation.
The question I was trying to answer is "How can we identify code that is statically reachable but is not needed at runtime (and we might be able to get rid of it by reorganizing the code a bit)?". The answer to that is profiling and profile-based code generation.
The idea is simple:
* We build a special version of the program that keeps track of what methods executed.
* At process termination, this data is written out to a file.
* We then recompile the program, passing in the method list as one of the inputs.
* When compiling a method that was in the list, compile as usual.
* When compiling a method that wasn’t in the list, replace it with a failfast.
* We still run the usual dependency analysis so the failfast methods are going to stop graph expansion at method boundaries. One could do better than this (cut off at basic block boundaries) but that’s a lot more work with a small benefit.
The profiling file format is simple: list of assembly MVIDs, followed by an array of bools. Index within the array corresponds to a MethodDef token within the assembly. So if bool at index N of assembly A is true, method in assembly with MVID A at RID N is reachable. Using MVIDs (GUID) instead of assembly name helps identify mismatches between profile data and input assemblies. Best to build with `-p:Deterministic=true`.
# Usage
Build the app with `-p:_InstrumentReachability=true`. Run the app and exercise all the necessary codepaths. This produces a reach.mprof file in the current directory. You can run the app multiple times, the profile data will get merged automatically.
Then build the app again with `-p:_ReachabilityInstrumentationFile=reach.mprof`. Output of this build will be a profile-based version.
Just to give an idea of how small things will be:
* dotnet new webapiaot: 4.1 MB (down from 8.7 MB)
* Hello World with OptimizationPreference=Size, StackTraceSupport=false, UseSystemResourceKeys=true: 425 kB
The app may or may not actually work. Hello world works fine. Webapiaot hits an issue where our ability to optimize things better in the profile-based version leads to new codepaths being executed. It can be worked around by forcing something that triggers the optimization into profile data. For example, for webapiaot this helps:
```csharp
typeof(Holder).GetCustomAttribute<MyAttribute>().TheType.IsConstructedGenericType.ToString();
class MyAttribute : Attribute { public MyAttribute(Type t) => TheType = t; public Type TheType; }
class Gen<T> { }
[My(typeof(Gen<>))]
class Holder { }
```
(We need to call `IsConstructedGenericType` on a type that got its MethodTable optimized away.)
It is not strictly necessary for the outputs to be runnable for this to be useful. The idea is to generate DGML/MSTAT of the unprofiled version, then DGML/MSTAT of the profiled version, and diff them. The diff might highlight parts of the app that could maybe be removed by reorganizing the code.
Not everything will be removable. For example, in the Hello World, most of exception handling is gone, but one can still run into exceptions at runtime (I just didn’t when I profiled it).
# Implementation
There are 3 parts: compiler component to generate instrumented outputs, CoreLib change to save profile data on app exist, and compiler component to consume profile data.
## Generating profile data
ReachabilityInstrumentationProvider is the main workhorse. It’s an ILProvider that wraps whatever IL we got from the input assembly and prefixes each method body with two instructions: ldc.i4.1 followed by stsfld. The stsfld targets a compiler-generated RVA static field. The compiler lays out all these fields in a way that their in-memory positions correspond to profile data positions (so saving profile data just means copying a memory range to a file).
ReachabilityDataBlobNode is responsible for creating the data blob an laying out all the RVA static fields that the code refers to.
Last but not least, InitializeMethod generates a small stub that informs corelib where to find the data blob at runtime. We hook up InitializeMethod into StartupCodeMain.
## Saving profile data
Within CoreLib, we define two methods - one is called at startup and informs us where profile data blob lives. The other is the very last managed method executed. It saves the data blob to a file.
## Consuming profile data
This is another ILProvider that either returns the underlying IL unmodified, or replaces it with a failfast call.
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 12, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-NativeAOT-coreclrHackathonIssues picked for Hackathon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@MichalStrehovsky@EgorBo@kekekeks@GerardSmit@agocke@am11