Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Change temporary entrypoints to be lazily allocated by davidwrighton · Pull Request #101580 · dotnet/runtime · GitHub
Skip to content

Change temporary entrypoints to be lazily allocated - #101580

Merged
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints
Jul 14, 2024
Merged

Change temporary entrypoints to be lazily allocated#101580
davidwrighton merged 59 commits into
dotnet:mainfrom
davidwrighton:change_temporary_entrypoints

Conversation

@davidwrighton

@davidwrightondavidwrighton commented Apr 25, 2024

Copy link
Copy Markdown
Member

This change moves temporary entrypoints to be allocated when they are needed instead of eagerly at type load time

As you can see, with a very small test case the wins are significant, and notably, reducing allocations in the stub heaps is nice as they are difficult to allocate efficiently as they require a bunch of memory mapping operations. In addition, with this change I'm removing the concept of Compact Entrypoints, as they are no longer significantly beneficial as of the advent of tiered compilation, and would be difficult to integrate with this change.

Baseline
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
PR
Loader Heap:
----------------------------------------
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1ad000 (1757184) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x38000 (229376) bytes total.
NewStubPrecodeHeap: Size: 0x8000 (32768) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x2df000 (3010560) bytes total, 0x3000 (12288) bytes wasted.

What can be seen here, is that the size of the Precode heaps shrinks dramatically, and the size of the HighFrequency heap grows a little bit. With a larger test case (launching powershell with a -c exit command line argument, the gains are still very similar. Also with powershell, I have a simple benchmark which runs pwsh.exe -c exit 100 times and measures the total execution time. The performance improvement is around 1% which I consider to be quite a nice result.

Architecture/VersionTotal SizeProcess Execution Time
Baseline X869,420,800 bytes23.33 seconds
PR X866,348,800 bytes23.18 seconds
Baseline X6413,586,432 bytes23.23 seconds
PR X6410,440,704 bytes23.15 seconds

Baseline
Loader Heap:
----------------------------------------
System Domain: 7ffab916ec00
LoaderAllocator: 7ffab916ec00
LowFrequencyHeap: Size: 0xf0000 (983040) bytes total.
HighFrequencyHeap: Size: 0x16a000 (1482752) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x168000 (1474560) bytes total.
NewStubPrecodeHeap: Size: 0x18000 (98304) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x3dd000 (4050944) bytes total, 0x3000 (12288) bytes wasted.
Compare
Loader Heap:
----------------------------------------
System Domain: 7ff9eb49dc00
LoaderAllocator: 7ff9eb49dc00
LowFrequencyHeap: Size: 0xef000 (978944) bytes total.
HighFrequencyHeap: Size: 0x1b2000 (1777664) bytes total, 0x3000 (12288) bytes wasted.
StubHeap: Size: 0x1000 (4096) bytes total.
FixupPrecodeHeap: Size: 0x70000 (458752) bytes total.
NewStubPrecodeHeap: Size: 0x10000 (65536) bytes total.
IndirectionCellHeap: Size: 0x1000 (4096) bytes total.
CacheEntryHeap: Size: 0x1000 (4096) bytes total.
Total size: Size: 0x324000 (3293184) bytes total, 0x3000 (12288) bytes wasted.
LowFrequencyHeap is 4KB bigger
HighFrequencyHeap is 288KB bigger
FixupPrecodeHeap is 992KB smaller
NewstubPrecodeHeap is 32KB smaller
@build-analysisbuild-analysisBot mentioned this pull request May 9, 2024
davidwrightonand others added 4 commits June 28, 2024 11:59
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Try removing EnsureSlotFilled
Implement IsEligibleForTieredCompilation in terms of IsEligibleForTieredCompilation_NoCheckMethodDescChunk
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/debug/daccess/request.cpp
Comment threadsrc/coreclr/vm/clsload.cpp
if (m_pRepresentativeMT != NULL)
{
pMD = m_pRepresentativeMT->GetMethodDescForSlot(m_representativeSlot);
pMD = m_pRepresentativeMT->GetMethodDescForSlot_NoThrow(m_representativeSlot);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My point is if this is not supposed to return a NULL and we're calling the "NoThrow" then at the very least we need an assert to convey that assumption.

Comment threadsrc/coreclr/vm/methodtable.h
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment threadsrc/coreclr/vm/method.cpp Outdated
Co-authored-by: Aaron Robinson <arobins@microsoft.com>
…can return NULL ... and then another thread could produce a stable entrypoint and the assert could lose the race
Comment threadsrc/coreclr/vm/method.cpp
Comment threadsrc/coreclr/vm/method.cpp Outdated
Comment on lines +3098 to +3105
if (RequiresStableEntryPoint())
{
GetOrCreatePrecode()->SetTargetInterlocked(entryPoint);
}
else
{
SetStableEntryPointInterlocked(entryPoint);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code reads odd. If RequiresStableEntryPoint() returns false, then we set the stable entry point? If this is correct, I think a comment is appropriate.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sigh. It is right, but yeah, this should be better structured. A RequiredStableEntryPoint is currently required have a stable entrypoint which is a Precode. Otherwise, we can generate set a stable entrypoint which isn't based on a Precode. Before this change, we would have reliably hit the HasPrecode flag above, but with making Precode generation lazy, this is a point where the Precode needs to be forced at this point. However, I think I'll fix this by changing the previous condition to be (RequiresStableEntryPoint()) and not use GetPrecode, and instead use GetOrCreatePrecode(). Then this portion of the condition can revert to be what it was before.

@davidwrighton
davidwrighton merged commit dacf9db into dotnet:mainJul 14, 2024
davidwrighton added a commit to dotnet/diagnostics that referenced this pull request Jul 17, 2024
- Support iterating methods from the new ISOSDacInterface15
- Re-plumb the existing logic as a local implementation of
ISOSDacInterface15 on top of apis that have been present for a long time
This is dependent on dotnet/runtime#101580
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 14, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@davidwrighton@jkotas@AaronRobinsonMSFT