Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); More precise writebarrier for regions by PeterSolMS · Pull Request #67389 · dotnet/runtime · GitHub
Skip to content

More precise writebarrier for regions - #67389

Merged
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions
Aug 9, 2022
Merged

More precise writebarrier for regions#67389
PeterSolMS merged 29 commits into
dotnet:mainfrom
PeterSolMS:More_precise_writebarrier_for_regions

Conversation

@PeterSolMS

Copy link
Copy Markdown
Contributor

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

…g 4 bits for the current and planned generation. WKS shows no overall improvement, SVR crashes.
…l limits where objects in ephemeral regions may be located. We have to do a range check on the child object in mark_through_cards_helper anyway, and using the ephemeral range allows us to skip the table lookup in the cases where a child object cannot possibly be in an ephemeral region.
- accidentally removed setting plan gen num
- need to make the default write barrier larger so we have enough space
- fix copy & paste issue in GetCurrentWriteBarrierCode
…ons would remain stuck. This was because in these cases we would not explore the complete range of card table entries for the card bundle.
 - use BitScanForward in find_card, find_card_dword
- when we start a new card dword, consult the card bundles first
- change JIT_ByRefWriteBarrier to consult the region_to_generation_table and set only single bits in the card table.
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Issue Details

This introduces a lookup table for regions where we can find the current generation and the planned generation efficiently.

The table has byte-sized elements where the low nibble is the current generation and the high nibble is the planned generation.

The table is used in mark_through_cards_helper and in the write barriers (for now only the most frequently used ones, Array.Copy has its own way of setting cards that I haven't fixed).

I have changed the write barrier to only set single bits for the case where a pointer to younger generation is stored into an object in an older generation. This costs an interlocked operation in the case the bit is not already set. Hopefully though this will be more than compensated by lower cost in card marking.

I haven't implemented yet committing only the part of the lookup table that is needed.

Author:PeterSolMS
Assignees:PeterSolMS
Labels:

area-GC-coreclr

Milestone:-

Comment threadsrc/coreclr/gc/gc.cpp Outdated
}
if (ephemeral_change)
{
stomp_write_barrier_ephemeral (ephemeral_low, ephemeral_high,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stomp_write_barrier_ephemeral is required to be called while the EE is suspended. if we are calling this from init_heap_segment, it means it can be called when a new gen0 region is acquired while the EE is running.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and when it's called from find_first_valid_region, we could optimize and only call this once at the end of the GC (if the ephemeral range actually got larger).

…f just a single bit. This will allows us to determine the tradeoff between being more precise in the write barrier, which saves work in card marking, and being faster in the write barrier which causes more work in card marking.
@Maoni0

Maoni0 commented Jul 7, 2022

Copy link
Copy Markdown
Member

running on a 1st party prod workload -

indexBaselineNewDiffDiff %
3Process Duration (Sec)53,887.6353,858.88-28.751-0.053
4Total Allocated MB16,252,981.5815,819,053.88-433,927.69-2.67
5Max Size Peak MB26,878.4427,017.04138.6070.516
6GC Count7,269.007,573.003044.182
7Heap Count484800
8Gen0 Count3,630.003,779.001494.105
9Gen1 Count3,466.003,621.001554.472
10Ephemeral Count7,096.007,400.003044.284
11Gen2 Blocking Count4400
12BGC Count16916900
13Gen0 Total Pause Time MSec269,123.47231,960.87-37,162.60-13.81
14Gen1 Total Pause Time MSec355,468.13337,579.93-17,888.20-5.032
15Ephemeral Total Pause Time MSec624,591.60569,540.80-55,050.80-8.814
16Blocking Gen2 Total Pause Time MSec4,260.202,376.57-1,883.63-44.22
17BGC Total Pause Time MSec14,630.2513,270.05-1,360.20-9.297
18GC Pause Time %1.1941.087-0.108-9.011
19Avg. Gen0 Pause Time (ms)74.13961.382-12.757-17.21
20Avg. Gen1 Pause Time (ms)102.55993.228-9.33-9.097
21Avg. Gen0 Promoted (mb)170.597166.073-4.524-2.652
22Avg. Gen1 Promoted (mb)343.586332.5-11.087-3.227
23Avg. Gen0 Speed (mb/ms)2.3012.7060.40517.58
24Avg. Gen1 Speed (mb/ms)3.353.5670.2166.458

looking at 500 GCs during steady state as an example -

image

…contain flags to indicate whether a region is sweep-in-plan, and whether it has been demoted.
Generalize the config setting to change the write barrier to allow reverting to the SVR type write barrier as well.
Bug fixes concerning setting the ephemeral limits in the write barrier, and where to compute the ephemeral limits within the GC.
Use the lookup via the map_region_to_generation table in the mark phase as well.
…dy relocated and thus shouldn't be tested against gc_low/gc_high.
…doesn't need to be updated between GCs.
Removed file name argument to _ASSERTE_ALL_BUILDS macro.
… Volatile<T> expands to T volatile.
Use fixed ephemeral bounds for now, but keep more sophisticated code for setting ephemeral_low around.
@AndyAyersMS

Copy link
Copy Markdown
Member

This improved crossgen2 throughput:
newplot - 2022-09-08T093600 033

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@PeterSolMS@Maoni0@AndyAyersMS