PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

PGO: Instrument cold blocks with profile validators - #53840

Closed
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator
Closed

PGO: Instrument cold blocks with profile validators#53840
EgorBo wants to merge 5 commits into
dotnet:mainfrom
EgorBo:pgo-validator

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 7, 2021

Copy link
Copy Markdown
Member

Profile Data can't be 100% reliable and there are cases where a hot or a semi-hot code ends up in a block with zero weight (aka "cold block"):

  • Static PGO that we ship is based on TE benchmarks and can be less relevant for other workloads.
  • We don't support context-sensitive profiling yet. It's not a big problem for static PGO where we don't promote methods to tier1 during profile collection, but it is a problem with the dynamic one where we can bake some weights during first seconds of the app session (it's enough to call a method 30 times) and use them for all other callsites. (see example below)
  • Mistakes in the logic where we scale/propagate/mix weights in the JIT.
  • Workflow can change over time and the blocks initially recognized as cold can become hot. The only way to fix this is the ability to deoptimize code and re-collect profile.
  • Sampling-based profiling is even less accurate if we switch to that.

In order to get a sense of the big picture, I'm introducing a sort of a late instrumentation where I insert helper calls into every cold block with a profile data at some late JIT Phase. I decided to use calls instead of counters to be able to quickly compose a CSV report on app shutdown. It allows me to quickly get a statistics which methods have problematic weights.

Example:

staticvoidMain(){// Promote DoWork to tier1 and bake weightsfor(inti=0;i<100;i++){DoWork(i);Thread.Sleep(16);}// Weights in DoWork are baked at this point.// Call it with unusual (from profile's point of view) weight.DoWork(-10);DoWork(-20);}[MethodImpl(MethodImplOptions.NoInlining)]staticintDoWork(inti){if(i>=0){return42;// always taken in the first loop}return-42;}

And run with:

DOTNET_TieredPGO=1
DOTNET_ProfileValidationPath="C:\prj\report.csv"

It will save a list of problematic methods where cold blocks were hit (in a CSV format) on app exit:
image

I tried to run a desktop app AvaloniaILSpy with default parameters and here is what it printed (I closed the form by hands after 10 seconds and random clicking):

image
(38 methods)

If I run it with DOTNET_TieredPGO=1 and DOTNET_TC_QuickJitForLoops=1 it lists 422 methods

PS: It doesn't tell which blocks specifically have invalid weights - it's just for overall sense.
PS2: I guess I should use full method names

/cc @AndyAyersMS@davidwrighton @dotnet/jit-contrib

@ghostghost added the area-VM-coreclr label Jun 7, 2021
@AndyAyersMS

Copy link
Copy Markdown
Member

Can you give some examples where this tech helped you sort out a problem?

I like the idea, but if we're going to go down this road I think we should look at something that has more long-lasting diagnostic capabilities, like doing a late instrumentation pass on all blocks, or relying on something like PIN.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Can you give some examples where this tech helped you sort out a problem?

I was just testing how much we can trust profile data (for inlining). Static PGO looks good so far in different benchmarks, there were some issues, like PowerShell benchmarks had in the hot path Path::HasExtension with different expectations.
Same for String::Equals, CastHelpers::ChkCast_Helper.

A small benchmark that just deserializes a complicated JSON file also had issues:

image

I can try to rewrite it to counters if you give me some pointers where to look at. But I guess it's not a priority for net6.0 so I'll just leave this as a prototype.

@EgorBoEgorBo closed this Jun 8, 2021
@ghostghost locked as resolved and limited conversation to collaborators Jul 8, 2021
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@AndyAyersMS