Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add ee_alloc_context - #104849

Merged
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit
Oct 15, 2024
Merged

Add ee_alloc_context#104849
noahfalk merged 7 commits into
dotnet:mainfrom
noahfalk:combined_limit

Conversation

@noahfalk

@noahfalknoahfalk commented Jul 13, 2024

Copy link
Copy Markdown
Member

This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.

A new field called ee_alloc_context::combined_limit should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.

There is another PR for NativeAOT making the same change: #104851

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

This is ready for review now.
All failures are known pre-existing failures according to build analysis. I believe I've applied your feedback as well @jkotas

@noahfalk
noahfalk marked this pull request as draft July 15, 2024 05:51
@noahfalk

Copy link
Copy Markdown
MemberAuthor

I converted back to draft because testing on NativeAOT revealed a race condition that I believe effects the CoreCLR work as well. I am investigating.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

All issues I know of are resolved and build analysis is green, so this is ready for review once more. Thanks!

Comment threadsrc/coreclr/debug/daccess/dacdbiimpl.cpp Outdated
Comment threadsrc/coreclr/debug/daccess/request.cpp Outdated
Comment threadsrc/native/managed/cdacreader/src/DataType.cs
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcheaputilities.h Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@jkotas

jkotas commented Jul 16, 2024

Copy link
Copy Markdown
Member

Is it that hard to implement it in a way that it does not introduce cache contention from thread enumerations running in parallel? I think it is preferable to go with designs that do not have performance issues by construction.

Comment threadsrc/coreclr/vm/gcenv.ee.cpp Outdated
@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas Following up on perf, I attempted to stress the threading by running gcperfsim with 400 threads and gc server config, allocating 500GB of objects. I used two different configs, one with a 50MB heap size and small objects and the other with a 2gb heap size allowing for a 20% LOH allocation mix to introduce more BGCs. My machine has 20 hardware threads. Median execution time in scenario 1 was 34.5 sec for PR and baseline. Scenario 2 was 38.4 sec for baseline and 38.5 sec for the PR. Noise on machine causes runs to vary by up to 1 second so the 0.1 diff isn't statistically significant. It's possible with a large sample size and a low noise environment we could discern some cost from the noise, but it certainly doesnt jump out. I also took CPU sampling traces which showed GCEnumAllocContexts consistently consumed 0.3% inclusive time before and after the change.

Also fyi I lost internet service at my house. I'm trying to get it restored but for now I only have GH access from my cell phone.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

@jkotas - anything remaining before you can sign off? I believe everything has been resolved aside from still waiting to hear from @elinor-fung or @lambdageek.

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I did some manual testing of the debugging scenario. In the normal DAC scenario I found and fixed one bug in the new commit. CDAC appeared to be working correctly. My test was running windbg + SOS with !threads and !DumpHeap to exercise reading gc_alloc_context. I confirmed using a debugger on the debugger that the CDAC case was executing down the CDAC branch of the DAC code as intended.

@elinor-fung

Copy link
Copy Markdown
Member

cDAC changes look good to me.

@jkotas

Copy link
Copy Markdown
Member

perf

To assess perf impact for changes like this, we typically try to craft a microbenchmark that tries to show worst regression or best improvement (one example: #96150 (comment)). Unless you introduce very bad bug, the impact will be lost in the noise from all other things going on in a general-purpose allocation throughput microbenchmark.

I am not worried about the perf impact too much with the latest version of the change, so I do not think it is necessary to do that.

@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-coreclr gcstress-extra

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

GC stress crash at:

 Assert failure(PID 2208 [0x000008a0], Thread: 9372 [0x249c]): Consistency check failed: AV in clr at this callstack:
------
CORECLR! Thread::CooperativeCleanup + 0xFF (0x6fff547e)
CORECLR! Thread::DetachThread + 0x124 (0x6fff69ec)
CORECLR! TlsDestructionMonitor::~TlsDestructionMonitor + 0x8C (0x7043d0d8)
CORECLR! _dyn_tls_dtor + 0x8A (0x7045116a)

It looks like a regression introduced by this change (Thread::CooperativeCleanup is modified by this change).

@noahfalk

Copy link
Copy Markdown
MemberAuthor

I've updated for the feedback so far + the test failure that showed up in the GC stress run. Looking at the time I think I may need to call it quits on this for .NET 9 and instead treat it as .NET 10 on a slower cadence. At this point I don't think I have more time available right now to do a staged checkin, revalidate testing after the ongoing changes, and continue to have time in reserve to respond to potential post-checkin issues.

noahfalkand others added 7 commits October 14, 2024 21:14
This change is some preparatory refactoring for the randomized allocation sampling feature. We need to add more state onto allocation context but we don't want to do a breaking change of the GC interface. The new state only needs to be visible to the EE but we want it physically near the existing alloc context state for good cache locality. To accomplish this we created a new ee_alloc_context struct which contains an instance of gc_alloc_context within it.
The new ee_alloc_context.combined_limit field should be used by fast allocation helpers to determine when to go down the slow path. Most of the time combined_limit has the same value as alloc_limit, but periodically we need to emit an allocation sampling event on an object that is somewhere in the middle of an AC. Using combined_limit rather than alloc_limit as the slow path trigger allows us to keep all the sampling event logic in the slow path.
combined_limit is now synchronized in GcEnumAllocContexts instead of RestartEE.
This requires the GC being constrained in how it updates the alloc_ptr and alloc_limit. No GC behavior changed,
in practice, but the constraints are now part of the EE<->GC contract so that we can rely on them in the EE code.
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
Co-authored-by: Jan Kotas <jkotas@microsoft.com>
The code of GetAllocContext() was constructing a PTR_gc_alloc_context which does a host->target pointer conversion. Those conversions work by doing a lookup in a dictionary of blocks of memory that we have previously marshalled and the pointer being converted is expected to be the start of the memory block. In this case we had never previously marshalled the gc_allocation_context on its own. We had only marshalled the m_pRuntimeThreadLocals block which includes the gc_allocation_context inside of it at a non-zero offset. This caused the host->target pointer conversion to fail which in turn meant commands like !threads in SOS would fail.
The fix is pretty trivial. We don't need to do a host->target conversion here at all because the calling code in the DAC is going to immediately convert right back to a host pointer. We can avoid the conversion in both directions by eliminating the cast and returning the host pointer directly.
Somehow a test reached Thread::CooperativeCleanup() with m_pRuntimeThreadLocals==NULL. Looking at the code I expected that would mean GetThread()==NULL or ThreadState contains TS_Dead, but neither of those conditions were true so it is unclear how it executed to that state. The callstack was:
>	coreclr.dll!Thread::CooperativeCleanup() Line 2762	C++
coreclr.dll!Thread::DetachThread(int fDLLThreadDetach) Line 936	C++
coreclr.dll!TlsDestructionMonitor::~TlsDestructionMonitor() Line 1745	C++
coreclr.dll!__dyn_tls_dtor(void * __formal, const unsigned long dwReason, void * __formal) Line 122	C++
ntdll.dll!_LdrxCallInitRoutine@16()	Unknown
ntdll.dll!LdrpCallInitRoutine()	Unknown
ntdll.dll!LdrpCallTlsInitializers()	Unknown
ntdll.dll!LdrShutdownThread()	Unknown
ntdll.dll!RtlExitUserThread()	Unknown
Regardless, the previous code wouldn't have hit that issue because it obtained the pointer through t_runtime_thread_locals rather than m_pRuntimeThreadLocals. I restored using t_runtime_thread_locals in the Cleanup routine. Out of caution I also searched for any other places that were previously accessing the alloc_context through the thread local address and ensured they don't switch to use m_pRuntimeThreadLocals either.
@noahfalk

noahfalk commented Oct 15, 2024

Copy link
Copy Markdown
MemberAuthor

I'm resuming work on this feature. @jkotas as best I'm aware all the comments you had made on the PR had been addressed and the GC stress issue was already fixed back in July. Let me know if there is anything else before we get this merged. Thanks!

Some cdac file moves had created merge conflicts so I've rebased and CI tests are running once again.

@noahfalk
noahfalk merged commit 69dc5ec into dotnet:mainOct 15, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Nov 15, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@noahfalk@jkotas@elinor-fung@VSadov