Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Remove List<T>.Enumerator.MoveNextRare - #118425

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2
Aug 6, 2025
Merged

Remove List<T>.Enumerator.MoveNextRare#118425
stephentoub merged 1 commit into
dotnet:mainfrom
stephentoub:listenumeratortake2

Conversation

@stephentoub

@stephentoubstephentoub commented Aug 6, 2025

Copy link
Copy Markdown
Member

Take 2 on #116150

The JIT work to stack allocate enumerators stops working with List<T> when List<T> gets sufficiently long, e.g. around 1000 elements. At that point, profiling sees the MoveNextRare method used for the last MoveNext as being cold and doesn't inline it. With it not inlined, the boxed enumerator escapes, and the enumerator is then not stack allocated.

(This change only moves the boundary significantly, from ~1000 elements to ~10000 elements. At ~10000, it starts hitting a new limit, due to OSR.)

CopilotAI review requested due to automatic review settings August 6, 2025 02:22
@github-actionsgithub-actionsBot added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Aug 6, 2025

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR optimizes the List<T>.Enumerator implementation by removing the MoveNextRare method to improve JIT stack allocation of enumerators. The change addresses a performance issue where the JIT's stack allocation optimization for enumerators fails when List<T> instances become sufficiently large (around 1000 elements) because the MoveNextRare method is considered cold and not inlined, causing the enumerator to be boxed instead of stack-allocated.

Key changes:

  • Inlines the MoveNextRare logic directly into the MoveNext method
  • Reorders field declarations and simplifies the enumerator state management
  • Changes the end-of-enumeration index marker from _list._size + 1 to -1

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoubstephentoub added this to the 10.0.0 milestone Aug 6, 2025
@stephentoubstephentoub added area-System.Collections and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Aug 6, 2025
@stephentoub

Copy link
Copy Markdown
MemberAuthor

Related to #118420, cc: @AndyAyersMS

@AndyAyersMS

Copy link
Copy Markdown
Member

FYI there is a similar pattern in PriorityQueue

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@EgorBot -arm -amd -intel

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Running;BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);[MemoryDiagnoser(false)]publicclassBench{[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumList(List<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}[Benchmark][ArgumentsSource(nameof(GetLists))]publicintSumEnumerable(IEnumerable<int>list){intsum=0;foreach(intiteminlist){sum+=item;}returnsum;}publicstaticIEnumerable<List<int>>GetLists()=>fromcountinnewint[]{1,10,1_000}selectEnumerable.Range(0,count).ToList();}

@stephentoub

Copy link
Copy Markdown
MemberAuthor

@jkotas, any concerns?

@jkotas

Copy link
Copy Markdown
Member

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

Should we go all the way and mark MoveNext() with aggressive inlining?

I do not have a strong opinion either way. This feels like a variant of the code size vs. microbenchmark perf trade off we have faced number of times.

@stephentoub

Copy link
Copy Markdown
MemberAuthor

I assume that this will regress perf for common cases without PGO (NAOT and probably Mono too) since MoveNext() hot path is not going to be inlined anymore. Is that correct?

@EgorBo or @AndyAyersMS can comment more authoritatively, but it seems like it's still inlineable even without PGO:
SharpLab

I can mark it with AggressiveInlining, though, if we want to be more sure it happens.

@jkotas

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

@stephentoub
stephentoub enabled auto-merge (squash) August 6, 2025 21:06
@AndyAyersMS

Copy link
Copy Markdown
Member

it seems like it's still inlineable even without PGO:

Cool, LGTM then.

In cases where the jit sees the struct enumerator in IL, it applies a fair number of inlining boosts even without PGO:

Invoking compiler for the inlinee method System.Collections.Generic.List`1+Enumerator[int]:MoveNext():bool:this :
...
multiplier in methods of struct increased to 3.
9 ldfld or stfld over arguments which are structs. Multiplier increased to 4.
Inline candidate has arg that feeds range check. Multiplier increased to 5.
Inline candidate is generic and caller is not. Multiplier increased to 7.
Inline candidate has 1 foldable branches. Multiplier increased to 11.
Inline has 1 foldable binary expressions. Multiplier increased to 13.
Inline candidate has an arg that feeds a constant test. Multiplier increased to 14.
Inline candidate callsite is in a loop. Multiplier increased to 17.
calleeNativeSizeEstimate=829
callsiteNativeSizeEstimate=85
benefit multiplier=17
threshold=1445
Native estimate for function size is within threshold for inlining 82.9 <= 144.5 (multiplier = 17)

@stephentoub
stephentoub merged commit 8b08265 into dotnet:mainAug 6, 2025
137 checks passed
@stephentoub
stephentoub deleted the listenumeratortake2 branch August 6, 2025 22:24
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Sep 6, 2025
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@stephentoub@AndyAyersMS@jkotas@EgorBo