fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix for largepages with agressive decommit logic - #126929

Merged
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages
Apr 16, 2026
Merged

fix for largepages with agressive decommit logic#126929
mangod9 merged 3 commits into
dotnet:mainfrom
mangod9:fix/gc-largepages

Conversation

@mangod9

Copy link
Copy Markdown
Member

clear decommitted memory in the largepages scenario. Fixes#126903

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes a GC heap-corruption scenario when GCLargePages is enabled and an induced Aggressive GC triggers “decommit” bookkeeping that doesn’t actually decommit at the OS level for large pages. The change ensures the memory that is treated as decommitted is explicitly cleared so stale references can’t be observed later.

Changes:

  • In the induced-aggressive path of gc_heap::distribute_free_regions, clear the region tail that would normally be decommitted/zeroed by the OS.
  • Gate the clearing to use_large_pages_p, since only large pages make virtual_decommit a no-op while still updating GC bookkeeping.

@janvorlijanvorli left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you!

@janvorli

Copy link
Copy Markdown
Member

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

@mangod9

Copy link
Copy Markdown
MemberAuthor

@mangod9 I believe this change should get in as is. But I wonder if it would be better to integrate the clearing of used part of the large page into the virtual_decommit (adding an "end of used data" argument) in the future so that we prevent similar issues to occur due to some changes in the GC. I also wonder if all the other usages of virtual_decommit are fine for large pages w.r.t. the fact the memory is not cleared.

yeah moved it centrally to virtual_decommit now. I have looked through other large_pages code flow and this looks to be the only case.

@mangod9

Copy link
Copy Markdown
MemberAuthor

/ba-g downloading artifacts is constantly stuck on macOS

@mangod9
mangod9 merged commit 830b6fe into dotnet:mainApr 16, 2026
109 of 113 checks passed
// observes leftover object references after the region is reused.
if (use_large_pages_p && (end_of_data != nullptr) && (end_of_data > address))
{
memclr ((uint8_t*)address, (uint8_t*)end_of_data - (uint8_t*)address);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other paths, the GC just takes keeps track of the fact that memory is dirty and clears it right before it is used for allocations again in gc_heap::adjust_limit_clr. Would it be a better option here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the fix was following the same pattern like this in decommit_region:

if (require_clearing_memory_p)
{
uint8_t* clear_end = use_large_pages_p ? heap_segment_used (region) : heap_segment_committed (region);
size_t clear_size = clear_end - page_start;
memclr (page_start, clear_size);
heap_segment_used (region) = heap_segment_mem (region);
dprintf(REGIONS_LOG, ("cleared region %p(%p-%p) (%zu bytes)",
region,
page_start,
clear_end,
clear_size));
}
else
{
heap_segment_committed (region) = heap_segment_mem (region);
}

where memclr clears the full region for large_pages. Similar cleanup was missing during aggressive decommitting of tail regions.

cshung added a commit to cshung/runtime that referenced this pull request Apr 23, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
cshung added a commit to cshung/runtime that referenced this pull request Apr 24, 2026
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR dotnet#126929 fixed the resulting stale data corruption
by adding memclr in virtual_decommit, but this approach has downsides:
the memory is never returned to the OS, yet we pay for the clearing and
produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in dotnet#126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already
handles large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter
and no-op ternary added by dotnet#126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fixdotnet#126903
janvorli pushed a commit that referenced this pull request Apr 28, 2026
…7290)
With large pages, VirtualDecommit is a no-op since large pages cannot be
partially decommitted. PR #126929 fixed the resulting stale data
corruption by adding memclr in virtual_decommit, but this approach has
downsides: the memory is never returned to the OS, yet we pay for the
clearing and produce misleading committed/used bookkeeping.
Instead, skip the decommit entirely for large pages:
1. distribute_free_regions: skip the aggressive tail-region decommit
(the committed-but-unallocated tail of in-use regions). This was the
path that caused the heap corruption in #126903.
2. decommit_heap_segment: skip the whole-segment decommit used for
segment hoarding and BGC segment deletion. Same class of issue:
committed/used are lowered but physical memory retains stale data.
3. decommit_region: bypass virtual_decommit and call
reduce_committed_bytes directly, since decommit_region already handles
large pages correctly by clearing memory itself.
4. virtual_decommit: add an assert that it is never called for heap
memory when large pages are on. This catches any future caller that
forgets to handle the large pages case. The end_of_data parameter and
no-op ternary added by #126929 are removed.
Add GCLargePages=2 mode that simulates large pages using small pages:
sets use_large_pages_p=true but reserves with normal pages and commits
everything upfront. This exercises all large page GC code paths without
requiring OS large page setup or privileges, enabling CI testing.
Fix#126903
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 16, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GC heap corruption with GCLargePages

4 participants

@mangod9@janvorli@jkotas