Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add NoInlining to ArrayPool Rent/Return - #129319

Merged
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool
Jul 17, 2026
Merged

Add NoInlining to ArrayPool Rent/Return#129319
jakobbotsch merged 2 commits into
dotnet:mainfrom
jakobbotsch:no-inline-arraypool

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented Jun 12, 2026

Copy link
Copy Markdown
Member

These methods result in inlining based on a lot of questionable heuristics. For example, the benchmark in #124285 went from 200 bytes in .NET 9 to 1567 bytes in .NET 10 because of inlining these two methods. Additionally Return has a tendency to appear in finally clauses, and we do not penalize inlining into these (even though we probably should), so this can also disable finally cloning.

As a surgical fix late in the cycle this just marks these methods NoInlining.

Fix#124285

These methods result in inlining based on a lot of questionable
heuristics. For example, the benchmark in dotnet#124285 went from 200 bytes in
.NET 9 to 1567 bytes in .NET 10 because of inlining these two methods.
Additionally `Return` has a tendency to appear in `finally` clauses, and
we do not penalize inlining into these (even though we probably should),
so this can also disable finally cloning.
As a surgical fix late in the cycle this just marks these methods
`NoInlining`.
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR applies [MethodImpl(MethodImplOptions.NoInlining)] to the Rent and Return overrides on the primary ArrayPool<T> implementations in System.Private.CoreLib, to prevent the JIT from inlining these relatively large methods into callers.

Changes:

  • Mark SharedArrayPool<T>.Rent / Return as NoInlining.
  • Mark ConfigurableArrayPool<T>.Rent / Return as NoInlining (and add the required System.Runtime.CompilerServices using).

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
src/libraries/System.Private.CoreLib/src/System/Buffers/SharedArrayPool.csAdds NoInlining to Rent / Return on the shared pool implementation.
src/libraries/System.Private.CoreLib/src/System/Buffers/ConfigurableArrayPool.csAdds NoInlining to Rent / Return on the configurable pool implementation and adds the needed using.

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM.

It'd be good to get some SPMI and/or perf numbers included here and showcasing the improvement.

The general downside of NoInlining is of course that the JIT will not attempt to observe the IL and infer any heuristics about the method in question, but that should be fine here.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot

usingBenchmarkDotNet.Attributes;usingBenchmarkDotNet.Jobs;usingSystem;usingSystem.Buffers;usingSystem.Runtime.InteropServices;namespaceRentSpanUnmanagedPerfTests;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark(Description="ArrayPool + Span - foreach (SLOW on .NET 10 and .NET 9)")]publicintArrayPool_Foreach(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with foreachvarsum=0;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="ArrayPool + Span - for loop (OK on .NET 10 and .NET 9)")]publicintArrayPool_ForLoop(){varpool=ArrayPool<int>.Shared;vararray=pool.Rent(Size);try{varspan=newSpan<int>(array,0,Size);// Writefor(vari=0;i<Size;i++){span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<span.Length;i++){sum+=span[i];}returnsum;}finally{pool.Return(array);}}[Benchmark(Description="RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)")]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with foreachvarsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan){sum+=value;}returnsum;}[Benchmark(Description="RefStruct + Span field - for loop (OK on .NET 10 and .NET 9)")]publicintRefStruct_ForLoop(){usingvarwrapper=newSpanWrapper(Size);// Writefor(vari=0;i<Size;i++){wrapper.Span[i]=i*2;}// Read with for loopvarsum=0;for(vari=0;i<wrapper.Span.Length;i++){sum+=wrapper.Span[i];}returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose(){ArrayPool<int>.Shared.Return(m_Array);}}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot -intel

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

On my machine the results look like:

MethodJobToolchainMeanErrorStdDevRatio
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-ZICBXN\main\corerun.exe749.4 ns4.62 ns3.61 ns1.00
'RefStruct + Span field - foreach (SLOW on .NET 10, FAST on .NET 9)'Job-QENHHP\pr\corerun.exe467.5 ns2.08 ns1.74 ns0.62

Oddly Egorbot does not see an improvement, and when I look at its disassembly it has a well-optimized inner loop of

2026-06-1715:24:04.417 G_M000_IG08: ;; offset=0x00E02026-06-1715:24:04.417addecx, dword ptr [r14+rax]2026-06-1715:24:04.417addrax,42026-06-1715:24:04.417decedx2026-06-1715:24:04.417jne SHORT G_M000_IG08

whereas my inner loop without this fix looks like

G_M000_IG08: ;; offset=0x00E0moveax, dword ptr [rbp-0x3C]addecx, dword ptr [r14+4*rax]moveax, dword ptr [rbp-0x3C]inceaxmov dword ptr [rbp-0x3C],eaxcmp dword ptr [rbp-0x3C],0x3E8jl SHORT G_M000_IG08

Need to investigate a bit further what's different in the bot's environment.

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

@EgorBot --envvars DOTNET_JitDisasm:RefStruct_Foreach

usingBenchmarkDotNet.Attributes;usingSystem;usingSystem.Buffers;publicclassArrayPoolForeachRepro{privateconstintSize=1000;[Benchmark]publicintRefStruct_Foreach(){usingvarwrapper=newSpanWrapper(Size);for(vari=0;i<Size;i++)wrapper.Span[i]=i*2;varsum=0;varspan=wrapper.Span;foreach(ref readonly varvalueinspan)sum+=value;returnsum;}privatereadonlyrefstructSpanWrapper{privatereadonlyint[]m_Array;publicreadonlySpan<int>Span;publicSpanWrapper(intlength){m_Array=ArrayPool<int>.Shared.Rent(length);Span=newSpan<int>(m_Array,0,length);}publicvoidDispose()=>ArrayPool<int>.Shared.Return(m_Array);}}

@jakobbotsch

jakobbotsch commented Jul 17, 2026

Copy link
Copy Markdown
MemberAuthor

Still not seeing the perf improvement on egorbot, but I have validated this locally. I do see the large inlining happening even in egorbot's version, though the inner loop is efficient even with it. I am not totally sure what the difference is, but since the customer also has the perf problem I think we can roll with it.

@jakobbotsch
jakobbotsch marked this pull request as ready for review July 17, 2026 10:24
CopilotAI review requested due to automatic review settings July 17, 2026 10:24
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated no new comments.

@tannergooding

Copy link
Copy Markdown
Member

@jakobbotsch given you see a large improvement and EgorBot at best sees a small 5ns regression (generally within noise given the overall runtime), I'd say we take this and look at official perf lab results

@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

/ba-g Timeouts

@jakobbotsch
jakobbotsch merged commit 2119b86 into dotnet:mainJul 17, 2026
136 of 143 checks passed
@jakobbotsch
jakobbotsch deleted the no-inline-arraypool branch July 17, 2026 17:07
@dotnet-milestone-botdotnet-milestone-botBot added this to the 11.0-preview7 milestone Jul 18, 2026
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 17, 2026
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[JIT] JIT regression in .NET 10: foreach over Span<T> field in ref struct triggers excessive inlining of ArrayPool internals

5 participants

@jakobbotsch@tannergooding@EgorBo@MihaZupan